German Language LLM Index
a PeerBench project

Individual experiment result

MuSR (DE)

GPT-5.6 Terra OpenAI

Accuracy reasoning off provider-internal
78.7%
Accuracy

Run details

Date
2026-08-13
Test cases
553
Median latency
7.62s
Total cost
$2.26
Avg. prompt tokens
1,445.7
Avg. answer tokens
382.9
Avg. reasoning tokens
Quantization
provider-internal

About this benchmark

German MuSR v1.1 defect-filtered benchmark: 553 questions (11 truncated team-allocation translations excluded), generate_until cot+; primary metric accuracy.

2–5 option multiple choice translated · German translation Source ↗

Subject breakdown

Subject Score
murder mystery 82.8%
object placements 75.0%
team allocation 75.3%

Provenance

Canonical model
openai/gpt-5.6-terra
Run ID
2026-08-13T09-55-11__openrouter__openai__gpt-5.6-terra__musr_de