German Language LLM Index
a PeerBench project

Individual experiment result

MMLU-ProX (DE)

GPT-5.5 OpenAI

Exact match reasoning on / locked provider-internal
87.2%
Exact match

Run details

Date
2026-07-01
Test cases
11,759
Median latency
29.48s
Total cost
$193.51
Avg. prompt tokens
1,594.7
Avg. answer tokens
145.5
Avg. reasoning tokens
137.3
Quantization
provider-internal

About this benchmark

German MMLU-ProX across 14 subjects; primary metric accuracy. Legacy rows have 11,759 items. Corrected v1.1 rows have 11,737 after a frozen, model-independent manifest excludes 20 duplicated-gold translation collisions, one independently reviewed ambiguous source item, and one independently confirmed notation-translation collapse; the row's item count identifies the definition.

10-option multiple choice translated · Professional translation Source ↗

Subject breakdown

Subject Score
biology 92.3%
business 91.1%
chemistry 89.8%
computer science 90.0%
economics 89.2%
engineering 83.3%
health 81.8%
history 77.7%
law 74.8%
math 94.6%
other 84.3%
philosophy 85.2%
physics 90.8%
psychology 86.3%

Provenance

Canonical model
openai/gpt-5.5
Run ID
2026-07-01T09-02-53__openai__codex__gpt-5.5-low__mmlu_prox_de