German Language LLM Index
a PeerBench project

Individual experiment result

MMLU-ProX (DE)

Qwen3.8 27B (FP8) Alibaba

Exact match reasoning off fp8
80.0%
Exact match

Run details

Date
2026-08-15
Test cases
11,737
Median latency
19.03s
Total cost
Avg. prompt tokens
1,689.5
Avg. answer tokens
1,210.1
Avg. reasoning tokens
Quantization
fp8

About this benchmark

German MMLU-ProX across 14 subjects; primary metric accuracy. Legacy rows have 11,759 items. Corrected v1.1 rows have 11,737 after a frozen, model-independent manifest excludes 20 duplicated-gold translation collisions, one independently reviewed ambiguous source item, and one independently confirmed notation-translation collapse; the row's item count identifies the definition.

10-option multiple choice translated · Professional translation Source ↗

Subject breakdown

Subject Score
biology 89.9%
business 85.9%
chemistry 87.7%
computer science 83.4%
economics 85.1%
engineering 74.2%
health 70.9%
history 69.2%
law 53.8%
math 93.4%
other 70.2%
philosophy 68.5%
physics 87.7%
psychology 81.8%

Provenance

Canonical model
qwen/qwen3.8-27b-fp8
Run ID
2026-08-15T16-51-00__openai__Qwen__Qwen3.8-27B-FP8__mmlu_prox_de