German Language LLM Index
a PeerBench project

Individual experiment result

INCLUDE (DE)

MiniMax M2.7 MiniMax

Exact match reasoning on / locked fp8
72.7%
Exact match

Run details

Date
Not recorded
Test cases
139
Median latency
Total cost
Avg. prompt tokens
Avg. answer tokens
356.6
Avg. reasoning tokens
816.9
Quantization
fp8

About this benchmark

INCLUDE-DE v1.2: 107 valid German questions after a blinded semantic audit excluded 32 duplicate, wrong-keyed, non-unique, corrupted, or underspecified source rows consistently; strict final-answer parsing; primary metric exact match.

4-option multiple choice native · Native German Source ↗

Provenance

Canonical model
minimax/minimax-m2.7
Run ID
2026-06-08T13-30-45__openrouter__minimax__minimax-m2.7__include_de