German Language LLM Index
a PeerBench project

Individual experiment result

MMLU-ProX (DE)

Qwen3.6 27B (BF16) Alibaba

Exact match reasoning off bf16
82.3%
Exact match

Run details

Date
2026-08-15
Test cases
11,737
Median latency
18.46s
Total cost
Avg. prompt tokens
1,689.5
Avg. answer tokens
932.8
Avg. reasoning tokens
Quantization
bf16

About this benchmark

German MMLU-ProX across 14 subjects; primary metric accuracy. Legacy rows have 11,759 items. Corrected v1.1 rows have 11,737 after a frozen, model-independent manifest excludes 20 duplicated-gold translation collisions, one independently reviewed ambiguous source item, and one independently confirmed notation-translation collapse; the row's item count identifies the definition.

10-option multiple choice translated · Professional translation Source ↗

Subject breakdown

Subject Score
biology 89.8%
business 87.7%
chemistry 90.2%
computer science 88.5%
economics 86.2%
engineering 77.3%
health 75.1%
history 72.6%
law 57.1%
math 93.7%
other 74.3%
philosophy 73.9%
physics 89.7%
psychology 82.5%

Provenance

Canonical model
qwen/qwen3.6-27b
Run ID
2026-08-15T17-46-46__openai__Qwen__Qwen3.6-27B__mmlu_prox_de