German Language LLM Index
a PeerBench project

Individual experiment result

MuSR (DE)

Qwen3.5 9B Alibaba

Accuracy reasoning off bf16
80.9%
Accuracy

Run details

Date
2026-06-11
Test cases
564
Median latency
61.14s
Total cost
$0.22
Avg. prompt tokens
1,469.2
Avg. answer tokens
1,629.7
Avg. reasoning tokens
Quantization
bf16

About this benchmark

German MuSR v1.1 defect-filtered benchmark: 553 questions (11 truncated team-allocation translations excluded), generate_until cot+; primary metric accuracy.

2–5 option multiple choice translated · German translation Source ↗

Subject breakdown

Subject Score
murder mystery 77.6%
object placements 73.4%
team allocation 86.0%

Provenance

Canonical model
qwen/qwen3.5-9b
Run ID
2026-06-11T14-33-49__openai__qwen35-9b-bf16__musr_de