Individual experiment result
MuSR (DE)
GPT-5.5 OpenAI
Accuracy reasoning on / locked provider-internal
86.3%
Accuracy
Run details
- Date
- 2026-06-13
- Test cases
- 564
- Median latency
- 13.58s
- Total cost
- $14.53
- Avg. prompt tokens
- 1,449.5
- Avg. answer tokens
- 414.8
- Avg. reasoning tokens
- 202.4
- Quantization
- provider-internal
About this benchmark
German MuSR v1.1 defect-filtered benchmark: 553 questions (11 truncated team-allocation translations excluded), generate_until cot+; primary metric accuracy.
Subject breakdown
| Subject | Score |
|---|---|
| murder mystery | 90.0% |
| object placements | 68.8% |
| team allocation | 87.2% |
Provenance
- Canonical model
- openai/gpt-5.5
- Run ID
2026-06-13T18-12-13__openai__codex__gpt-5.5-low__musr_de