German Language LLM Index
a PeerBench project

Individual experiment result

MMLU-ProX (DE)

Claude Haiku 4.5 Anthropic

Exact match reasoning off provider-internal
75.3%
Exact match

Run details

Date
2026-06-03
Test cases
11,759
Median latency
3.76s
Total cost
$52.10
Avg. prompt tokens
2,262
Avg. answer tokens
433.7
Avg. reasoning tokens
Quantization
provider-internal

About this benchmark

German MMLU-ProX across 14 subjects; primary metric accuracy. Legacy rows have 11,759 items. Corrected v1.1 rows have 11,737 after a frozen, model-independent manifest excludes 20 duplicated-gold translation collisions, one independently reviewed ambiguous source item, and one independently confirmed notation-translation collapse; the row's item count identifies the definition.

10-option multiple choice translated · Professional translation Source ↗

Subject breakdown

Subject Score
biology 85.9%
business 79.0%
chemistry 80.4%
computer science 83.4%
economics 81.5%
engineering 64.7%
health 72.9%
history 65.4%
law 47.4%
math 86.8%
other 69.5%
philosophy 67.9%
physics 81.3%
psychology 78.9%

Provenance

Run ID
2026-06-03T15-52-51__openai__cc__claude-haiku-4-5-20251001__mmlu_prox_de