German Language LLM Index
a PeerBench project

Individual experiment result

MMLU-ProX (DE)

Gemma 4 12B Google

Exact match reasoning off bf16
73.1%
Exact match

Run details

Date
2026-06-11
Test cases
11,759
Median latency
92.81s
Total cost
Avg. prompt tokens
1,705.9
Avg. answer tokens
493.3
Avg. reasoning tokens
Quantization
bf16

About this benchmark

German MMLU-ProX across 14 subjects; primary metric accuracy. Legacy rows have 11,759 items. Corrected v1.1 rows have 11,737 after a frozen, model-independent manifest excludes 20 duplicated-gold translation collisions, one independently reviewed ambiguous source item, and one independently confirmed notation-translation collapse; the row's item count identifies the definition.

10-option multiple choice translated · Professional translation Source ↗

Subject breakdown

Subject Score
nothink biology 87.0%
nothink business 79.8%
nothink chemistry 79.9%
nothink computer science 79.3%
nothink economics 78.7%
nothink engineering 65.1%
nothink health 64.5%
nothink history 57.2%
nothink law 42.8%
nothink math 90.1%
nothink other 60.7%
nothink philosophy 60.3%
nothink physics 81.9%
nothink psychology 76.3%

Provenance

Canonical model
google/gemma-4-12b-it
Run ID
2026-06-11T15-08-57__openai__gemma4-12b-bf16__mmlu_prox_de_nothink