German Language LLM Index
a PeerBench project

Individual experiment result

MMLU-ProX (DE)

Gemini 3.1 Pro Google

Exact match reasoning on / locked provider-internal
89.2%
Exact match

Run details

Date
2026-08-15
Test cases
11,759
Median latency
3.46s
Total cost
$95.16
Avg. prompt tokens
1,668.9
Avg. answer tokens
243
Avg. reasoning tokens
153.3
Quantization
provider-internal

About this benchmark

German MMLU-ProX across 14 subjects; primary metric accuracy. Legacy rows have 11,759 items. Corrected v1.1 rows have 11,737 after a frozen, model-independent manifest excludes 20 duplicated-gold translation collisions, one independently reviewed ambiguous source item, and one independently confirmed notation-translation collapse; the row's item count identifies the definition.

10-option multiple choice translated · Professional translation Source ↗

Subject breakdown

Subject Score
nothink biology 95.1%
nothink business 91.8%
nothink chemistry 91.0%
nothink computer science 91.5%
nothink economics 91.4%
nothink engineering 86.0%
nothink health 82.8%
nothink history 82.7%
nothink law 78.6%
nothink math 95.0%
nothink other 85.1%
nothink philosophy 86.0%
nothink physics 93.6%
nothink psychology 90.2%

Provenance

Run ID
2026-08-15T13-46-32__vertex_ai__gemini-3.1-pro-preview__mmlu_prox_de_nothink