German Language LLM Index
a PeerBench project

Individual experiment result

MMLU-ProX (DE)

Qwen3.8 27B (BF16) Alibaba

Exact match reasoning off bf16
79.9%
Exact match

Run details

Date
2026-08-15
Test cases
11,737
Median latency
25.21s
Total cost
Avg. prompt tokens
1,689.5
Avg. answer tokens
1,198.3
Avg. reasoning tokens
Quantization
bf16

About this benchmark

German MMLU-ProX across 14 subjects; primary metric accuracy. Legacy rows have 11,759 items. Corrected v1.1 rows have 11,737 after a frozen, model-independent manifest excludes 20 duplicated-gold translation collisions, one independently reviewed ambiguous source item, and one independently confirmed notation-translation collapse; the row's item count identifies the definition.

10-option multiple choice translated · Professional translation Source ↗

Subject breakdown

Subject Score
biology 90.2%
business 85.1%
chemistry 88.3%
computer science 84.1%
economics 86.5%
engineering 73.2%
health 71.8%
history 70.5%
law 52.2%
math 93.3%
other 70.5%
philosophy 66.9%
physics 88.7%
psychology 79.5%

Provenance

Canonical model
qwen/qwen3.8-27b
Run ID
2026-08-15T08-21-28__openai__Qwen__Qwen3.8-27B__mmlu_prox_de