German Language LLM Index
a PeerBench project

Individual experiment result

MuSR (DE)

GPT-5.6 Luna OpenAI

Accuracy reasoning off provider-internal
75.8%
Accuracy

Run details

Date
2026-08-13
Test cases
553
Median latency
3.75s
Total cost
$0.21
Avg. prompt tokens
1,445.7
Avg. answer tokens
322.2
Avg. reasoning tokens
Quantization
provider-internal

About this benchmark

German MuSR v1.1 defect-filtered benchmark: 553 questions (11 truncated team-allocation translations excluded), generate_until cot+; primary metric accuracy.

2–5 option multiple choice translated · German translation Source ↗

Subject breakdown

Subject Score
murder mystery 80.4%
object placements 82.8%
team allocation 69.0%

Provenance

Canonical model
openai/gpt-5.6-luna
Run ID
2026-08-13T13-59-20__openrouter__openai__gpt-5.6-luna__musr_de