German Language LLM Index
a PeerBench project

Individual experiment result

MMLU-ProX (DE)

DeepSeek V4 Pro DeepSeek

Exact match reasoning off provider-internal
80.8%
Exact match

Run details

Date
2026-06-16
Test cases
11,759
Median latency
3.00s
Total cost
$3.71
Avg. prompt tokens
288.1
Avg. answer tokens
218.9
Avg. reasoning tokens
Quantization
provider-internal

About this benchmark

German MMLU-ProX across 14 subjects; primary metric accuracy. Legacy rows have 11,759 items. Corrected v1.1 rows have 11,737 after a frozen, model-independent manifest excludes 20 duplicated-gold translation collisions, one independently reviewed ambiguous source item, and one independently confirmed notation-translation collapse; the row's item count identifies the definition.

10-option multiple choice translated · Professional translation Source ↗

Subject breakdown

Subject Score
biology 90.2%
business 85.3%
chemistry 86.7%
computer science 85.4%
economics 84.1%
engineering 74.2%
health 74.2%
history 71.9%
law 55.3%
math 90.9%
other 77.3%
philosophy 73.7%
physics 86.8%
psychology 83.3%

Provenance

Run ID
2026-06-16T13-49-10__openrouter__deepseek__deepseek-v4-pro__mmlu_prox_de