German Language LLM Index
a PeerBench project

Individual experiment result

MMLU-ProX (DE)

Qwen3 14B Alibaba

Exact match reasoning off bf16
64.1%
Exact match

Run details

Date
2026-06-11
Test cases
11,759
Median latency
29.26s
Total cost
Avg. prompt tokens
1,887.3
Avg. answer tokens
413.7
Avg. reasoning tokens
Quantization
bf16

About this benchmark

German MMLU-ProX across 14 subjects; primary metric accuracy. Legacy rows have 11,759 items. Corrected v1.1 rows have 11,737 after a frozen, model-independent manifest excludes 20 duplicated-gold translation collisions, one independently reviewed ambiguous source item, and one independently confirmed notation-translation collapse; the row's item count identifies the definition.

10-option multiple choice translated · Professional translation Source ↗

Subject breakdown

Subject Score
nothink biology 81.3%
nothink business 70.3%
nothink chemistry 72.2%
nothink computer science 71.5%
nothink economics 71.1%
nothink engineering 55.9%
nothink health 56.8%
nothink history 48.8%
nothink law 27.4%
nothink math 82.2%
nothink other 52.1%
nothink philosophy 49.3%
nothink physics 73.1%
nothink psychology 66.0%

Provenance

Canonical model
qwen/qwen3-14b
Run ID
2026-06-11T15-03-02__openai__qwen3-14b-bf16__mmlu_prox_de_nothink