German Language LLM Index
a PeerBench project

Individual experiment result

MMLU-ProX (DE)

Qwen3.6 35B-A3B Alibaba

Exact match reasoning off fp8
80.0%
Exact match

Run details

Date
2026-05-31
Test cases
11,759
Median latency
4.67s
Total cost
$13.36
Avg. prompt tokens
1,689.6
Avg. answer tokens
984
Avg. reasoning tokens
Quantization
fp8

About this benchmark

German MMLU-ProX across 14 subjects; primary metric accuracy. Legacy rows have 11,759 items. Corrected v1.1 rows have 11,737 after a frozen, model-independent manifest excludes 20 duplicated-gold translation collisions, one independently reviewed ambiguous source item, and one independently confirmed notation-translation collapse; the row's item count identifies the definition.

10-option multiple choice translated · Professional translation Source ↗

Subject breakdown

Subject Score
biology 89.7%
business 83.5%
chemistry 88.2%
computer science 85.1%
economics 85.3%
engineering 77.5%
health 72.8%
history 69.0%
law 52.8%
math 92.7%
other 70.8%
philosophy 69.5%
physics 88.0%
psychology 78.7%

Provenance

Canonical model
qwen/qwen3.6-35b-a3b
Run ID
2026-05-31T12-32-33__openrouter__qwen__qwen3.6-35b-a3b