Individual experiment result
ScaLA (DE)
GLM-5.1 Z.ai
Matthews correlation coefficient reasoning off fp8
69.9%
Matthews correlation coefficient
Run details
- Date
- 2026-06-09
- Test cases
- 2,048
- Median latency
- 1.69s
- Total cost
- $2.57
- Avg. prompt tokens
- 901.6
- Avg. answer tokens
- 2.5
- Avg. reasoning tokens
- —
- Quantization
- fp8
About this benchmark
Native German linguistic acceptability — is the sentence grammatical (ja/nein)? generate_until; primary metric MCC (accuracy recorded as secondary).
Provenance
- Canonical model
- z-ai/glm-5.1
- Run ID
2026-06-09T15-44-40__openrouter__z-ai__glm-5.1__scala_de