German Language LLM Index
a PeerBench project

Individual experiment result

MMLU-ProX (DE)

Gemma 4 31B Google

Exact match reasoning off bf16
82.1%
Exact match

Run details

Date
2026-06-11
Test cases
11,759
Median latency
64.35s
Total cost
$4.53
Avg. prompt tokens
1,702.9
Avg. answer tokens
502.5
Avg. reasoning tokens
Quantization
bf16

About this benchmark

German MMLU-ProX across 14 subjects; primary metric accuracy. Legacy rows have 11,759 items. Corrected v1.1 rows have 11,737 after a frozen, model-independent manifest excludes 20 duplicated-gold translation collisions, one independently reviewed ambiguous source item, and one independently confirmed notation-translation collapse; the row's item count identifies the definition.

10-option multiple choice translated · Professional translation Source ↗

Subject breakdown

Subject Score
biology 91.2%
business 87.5%
chemistry 86.7%
computer science 85.9%
economics 86.8%
engineering 77.1%
health 74.1%
history 74.5%
law 59.0%
math 93.5%
other 73.7%
philosophy 74.3%
physics 89.0%
psychology 83.6%

Provenance

Canonical model
google/gemma-4-31b-it
Run ID
2026-06-11T15-10-23__openai__gemma-4-31b-it-bf16__mmlu_prox_de