German Language LLM Index
a PeerBench project

Individual experiment result

MMLU-ProX (DE)

Qwen3.5 9B Alibaba

Exact match reasoning off bf16
73.4%
Exact match

Run details

Date
2026-06-11
Test cases
11,759
Median latency
40.41s
Total cost
$3.63
Avg. prompt tokens
1,693.6
Avg. answer tokens
931.1
Avg. reasoning tokens
Quantization
bf16

About this benchmark

German MMLU-ProX across 14 subjects; primary metric accuracy. Legacy rows have 11,759 items. Corrected v1.1 rows have 11,737 after a frozen, model-independent manifest excludes 20 duplicated-gold translation collisions, one independently reviewed ambiguous source item, and one independently confirmed notation-translation collapse; the row's item count identifies the definition.

10-option multiple choice translated · Professional translation Source ↗

Subject breakdown

Subject Score
nothink biology 85.6%
nothink business 78.1%
nothink chemistry 85.0%
nothink computer science 78.0%
nothink economics 78.9%
nothink engineering 66.6%
nothink health 65.9%
nothink history 60.4%
nothink law 36.7%
nothink math 90.2%
nothink other 61.6%
nothink philosophy 58.1%
nothink physics 84.7%
nothink psychology 74.2%

Provenance

Canonical model
qwen/qwen3.5-9b
Run ID
2026-06-11T14-43-11__openai__qwen35-9b-bf16__mmlu_prox_de_nothink