German Language LLM Index
a PeerBench project

Individual experiment result

MuSR (DE)

DeepSeek V4 Pro DeepSeek

Accuracy reasoning off provider-internal
85.3%
Accuracy

Run details

Date
2026-06-23
Test cases
564
Median latency
12.60s
Total cost
$0.74
Avg. prompt tokens
1,657.8
Avg. answer tokens
719.7
Avg. reasoning tokens
Quantization
provider-internal

About this benchmark

German MuSR v1.1 defect-filtered benchmark: 553 questions (11 truncated team-allocation translations excluded), generate_until cot+; primary metric accuracy.

2–5 option multiple choice translated · German translation Source ↗

Subject breakdown

Subject Score
murder mystery 86.4%
object placements 70.3%
team allocation 88.0%

Provenance

Run ID
2026-06-23T13-55-55__openrouter__deepseek__deepseek-v4-pro__musr_de