German Language LLM Index
a PeerBench project

Individual experiment result

MuSR (DE)

Gemma 4 12B Google

Accuracy reasoning off bf16
81.6%
Accuracy

Run details

Date
2026-06-11
Test cases
564
Median latency
195.50s
Total cost
Avg. prompt tokens
1,477.9
Avg. answer tokens
707.4
Avg. reasoning tokens
Quantization
bf16

About this benchmark

German MuSR v1.1 defect-filtered benchmark: 553 questions (11 truncated team-allocation translations excluded), generate_until cot+; primary metric accuracy.

2–5 option multiple choice translated · German translation Source ↗

Subject breakdown

Subject Score
murder mystery 85.6%
object placements 79.7%
team allocation 78.0%

Provenance

Canonical model
google/gemma-4-12b-it
Run ID
2026-06-11T14-52-24__openai__gemma4-12b-bf16__musr_de