German Language LLM Index
a PeerBench project

Individual experiment result

MuSR (DE)

DeepSeek V4 Pro 0813 DeepSeek

Accuracy reasoning off provider-internal
82.3%
Accuracy

Run details

Date
2026-08-13
Test cases
553
Median latency
9.50s
Total cost
$0.67
Avg. prompt tokens
1,682.4
Avg. answer tokens
580.3
Avg. reasoning tokens
Quantization
provider-internal

About this benchmark

German MuSR v1.1 defect-filtered benchmark: 553 questions (11 truncated team-allocation translations excluded), generate_until cot+; primary metric accuracy.

2–5 option multiple choice translated · German translation Source ↗

Subject breakdown

Subject Score
murder mystery 85.6%
object placements 71.9%
team allocation 81.6%

Provenance

Run ID
2026-08-13T09-28-54__openrouter__deepseek__deepseek-v4-pro-0813__musr_de