Gemini 3.8 Flash
Highest German aggregate score
Independent German-language benchmarks for frontier and fast models — quality, latency, and cost, ranked for real deployment decisions.
Recommended picks
Highest German aggregate score
Strongest open-weight model in the ranking
High German score with measured fast throughput
Most German score per listed token price
No coding, vision, multimodal, or generic English arena score in the public aggregate.
Coverage is shown per model; partial averages are marked instead of hidden.
INCLUDE, GermEval, SB10K, ScaLA, MMLU-ProX, MMMLU, and MuSR.
Published rows are gated before they reach the static site.
Official German-language score from German benchmark evidence only. Partial coverage is shown in every row.
Showing 36 models that ran at least 3 of 6 German-language benchmarks (11 excluded for thin coverage). Cost is effective benchmark cost per 1,000 questions; speed is output tokens/sec.
Native German named-entity recognition — identify persons, locations, organisations and misc entities in German text, emitted as JSON. Scored with seqeval micro-F1 excluding the noisy MISC class. Run reasoning-off.
Native German exam and licensing questions covering region-specific knowledge — history, law, civics and culture. Written by humans in German, not translated.
Hard academic questions across 14 subjects — STEM, law, health, economics, philosophy and more. Professionally translated to German, with up to ten answer options per question.
| # | Model | Score | Latency | Tok in | Answer tok | Cost | Date |
|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Pro 🔒 reasoning | 89.2% | 3.46s | 1,668.9 | 243 | $95.16 | 2026-08-15 View result → |
| 2 | Gemini 3.8 Flash 🔒 reasoning | 89.0% | 7.24s | 1,664.9 | 222.4 | $104.01 | 2026-09-04 View result → |
| 3 | Gemini 3.7 Flash 🔒 reasoning | 88.2% | 4.60s | 1,664.9 | 191.5 | $34.62 | 2026-08-14 View result → |
| 4 | GPT-5.5 🔒 reasoning | 87.2% | 29.48s | 1,594.7 | 145.5 | $193.51 | 2026-07-01 View result → |
| 5 | Gemini 3.5 Flash | 86.5% | 2.23s | 1,664.9 | 353.3 | $66.76 | 2026-05-26 View result → |
| 6 | Gemini 3 Flash Preview | 86.3% | 3.26s | 1,664.9 | 311.5 | $20.78 | 2026-06-25 View result → |
| 7 | Gemini 3.6 Flash | 86.0% | 1.84s | 1,664.9 | 259.2 | $52.23 | 2026-07-23 View result → |
| 8 | Qwen3.6 27B (BF16) | 82.3% | 18.46s | 1,689.5 | 932.8 | — | 2026-08-15 View result → |
| 9 | Gemini 3.1 Flash-Lite | 82.2% | 1.24s | 1,666.1 | 304.5 | $10.21 | 2026-06-03 View result → |
| 10 | Gemma 4 31B | 82.1% | 64.35s | 1,702.9 | 502.5 | $4.53 | 2026-06-11 View result → |
| 11 | DeepSeek V4 Pro | 80.8% | 3.00s | 288.1 | 218.9 | $3.71 | 2026-06-16 View result → |
| 12 | DeepSeek V4 Pro 0813 | 80.7% | 2.92s | 1,718.4 | 184.8 | $3.44 | 2026-08-14 View result → |
| 13 | Gemini 3.5 Flash-Lite | 80.2% | 1.57s | 1,664.9 | 307.2 | $14.90 | 2026-07-23 View result → |
| 14 | Qwen3.6 35B-A3B | 80.0% | 4.67s | 1,689.6 | 984 | $13.36 | 2026-05-31 View result → |
| 15 | Qwen3.8 27B (FP8) | 80.0% | 19.03s | 1,689.5 | 1,210.1 | — | 2026-08-15 View result → |
| 16 | Qwen3.8 27B (BF16) | 79.9% | 25.21s | 1,689.5 | 1,198.3 | — | 2026-08-15 View result → |
| 17 | Gemini 2.5 Flash | 79.6% | 2.69s | 1,664.9 | 786 | $28.89 | 2026-05-26 View result → |
| 18 | Gemma 4 26B A4B | 78.2% | 7.77s | 1,700.6 | 662.1 | $4.64 | 2026-06-05 View result → |
| 19 | Claude Haiku 4.5 | 75.3% | 3.76s | 2,262 | 433.7 | $52.10 | 2026-06-03 View result → |
| 20 | DeepSeek V4 Flash | 74.9% | 1.98s | 1,718.7 | 162.2 | $1.04 | 2026-06-16 View result → |
| 21 | Qwen3.5 9B | 73.4% | 40.41s | 1,693.6 | 931.1 | $3.63 | 2026-06-11 View result → |
| 22 | Tencent HY3-Preview | 73.2% | 6.14s | 1,910.2 | 586.3 | $2.76 | 2026-05-27 View result → |
| 23 | Gemma 4 12B | 73.1% | 92.81s | 1,705.9 | 493.3 | — | 2026-06-11 View result → |
| 24 | Gemini 2.5 Flash-Lite | 71.2% | 2.48s | 1,665.1 | 1,493.5 | $8.95 | 2026-06-03 View result → |
| 25 | Qwen3 14B | 64.1% | 29.26s | 1,887.3 | 413.7 | — | 2026-06-11 View result → |
| Subject | Model | Score |
|---|---|---|
| biology | Gemini 3.7 Flash | 94.0% |
| biology | Gemini 3.8 Flash | 93.7% |
| biology | Gemini 3.5 Flash | 92.9% |
| biology | Gemini 3 Flash Preview | 92.6% |
| biology | Gemini 3.6 Flash | 92.6% |
| biology | GPT-5.5 | 92.3% |
| biology | Gemma 4 31B | 91.2% |
| biology | DeepSeek V4 Pro | 90.2% |
| biology | Qwen3.8 27B (BF16) | 90.2% |
| biology | DeepSeek V4 Pro 0813 | 90.1% |
| biology | Qwen3.8 27B (FP8) | 89.9% |
| biology | Qwen3.6 27B (BF16) | 89.8% |
| biology | Qwen3.6 35B-A3B | 89.7% |
| biology | Gemini 3.1 Flash-Lite | 89.4% |
| biology | Gemini 2.5 Flash | 88.8% |
| biology | DeepSeek V4 Flash | 88.4% |
| biology | Gemini 3.5 Flash-Lite | 88.1% |
| biology | Gemma 4 26B A4B | 87.3% |
| biology | Gemini 2.5 Flash-Lite | 86.8% |
| biology | Claude Haiku 4.5 | 85.9% |
| biology | Tencent HY3-Preview | 85.8% |
| business | Gemini 3.8 Flash | 92.3% |
| business | GPT-5.5 | 91.1% |
| business | Gemini 3.7 Flash | 90.6% |
| business | Gemini 3.5 Flash | 89.6% |
| business | Gemini 3 Flash Preview | 89.5% |
| business | Gemini 3.6 Flash | 88.2% |
| business | Qwen3.6 27B (BF16) | 87.7% |
| business | Gemma 4 31B | 87.5% |
| business | Qwen3.8 27B (FP8) | 85.9% |
| business | Gemini 3.1 Flash-Lite | 85.3% |
| business | DeepSeek V4 Pro | 85.3% |
| business | Qwen3.8 27B (BF16) | 85.1% |
| business | Gemini 3.5 Flash-Lite | 84.4% |
| business | Gemma 4 26B A4B | 83.9% |
| business | Gemini 2.5 Flash | 83.5% |
| business | Qwen3.6 35B-A3B | 83.5% |
| business | DeepSeek V4 Pro 0813 | 83.4% |
| business | Claude Haiku 4.5 | 79.0% |
| business | Tencent HY3-Preview | 78.3% |
| business | Gemini 2.5 Flash-Lite | 77.7% |
| business | DeepSeek V4 Flash | 76.2% |
| chemistry | Gemini 3.8 Flash | 91.8% |
| chemistry | Gemini 3.7 Flash | 91.4% |
| chemistry | Qwen3.6 27B (BF16) | 90.2% |
| chemistry | GPT-5.5 | 89.8% |
| chemistry | Gemini 3 Flash Preview | 89.2% |
| chemistry | Gemini 3.5 Flash | 89.1% |
| chemistry | Gemini 3.6 Flash | 88.3% |
| chemistry | Qwen3.8 27B (BF16) | 88.3% |
| chemistry | Qwen3.6 35B-A3B | 88.2% |
| chemistry | Qwen3.8 27B (FP8) | 87.7% |
| chemistry | Gemini 3.1 Flash-Lite | 87.1% |
| chemistry | Gemini 2.5 Flash | 86.9% |
| chemistry | Gemma 4 31B | 86.7% |
| chemistry | DeepSeek V4 Pro | 86.7% |
| chemistry | DeepSeek V4 Pro 0813 | 85.9% |
| chemistry | Gemma 4 26B A4B | 85.4% |
| chemistry | Gemini 3.5 Flash-Lite | 83.8% |
| chemistry | Claude Haiku 4.5 | 80.4% |
| chemistry | Gemini 2.5 Flash-Lite | 78.4% |
| chemistry | Tencent HY3-Preview | 75.6% |
| chemistry | DeepSeek V4 Flash | 73.3% |
| computer science | Gemini 3.8 Flash | 91.5% |
| computer science | Gemini 3.7 Flash | 90.7% |
| computer science | GPT-5.5 | 90.0% |
| computer science | Gemini 3.1 Flash-Lite | 88.8% |
| computer science | Gemini 3.6 Flash | 88.5% |
| computer science | Qwen3.6 27B (BF16) | 88.5% |
| computer science | Gemini 3.5 Flash | 87.6% |
| computer science | Gemini 3 Flash Preview | 86.6% |
| computer science | DeepSeek V4 Flash | 85.9% |
| computer science | Gemma 4 31B | 85.9% |
| computer science | DeepSeek V4 Pro | 85.4% |
| computer science | Gemini 2.5 Flash | 85.1% |
| computer science | Qwen3.6 35B-A3B | 85.1% |
| computer science | Qwen3.8 27B (BF16) | 84.1% |
| computer science | Gemma 4 26B A4B | 83.9% |
| computer science | DeepSeek V4 Pro 0813 | 83.6% |
| computer science | Claude Haiku 4.5 | 83.4% |
| computer science | Qwen3.8 27B (FP8) | 83.4% |
| computer science | Gemini 3.5 Flash-Lite | 83.2% |
| computer science | Gemini 2.5 Flash-Lite | 77.8% |
| computer science | Tencent HY3-Preview | 76.8% |
| economics | Gemini 3.8 Flash | 91.1% |
| economics | Gemini 3.7 Flash | 89.9% |
| economics | GPT-5.5 | 89.2% |
| economics | Gemini 3.5 Flash | 89.1% |
| economics | Gemini 3 Flash Preview | 88.7% |
| economics | Gemini 3.6 Flash | 88.7% |
| economics | Gemini 3.1 Flash-Lite | 87.3% |
| economics | Gemma 4 31B | 86.8% |
| economics | Qwen3.8 27B (BF16) | 86.5% |
| economics | Gemini 2.5 Flash | 86.4% |
| economics | Qwen3.6 27B (BF16) | 86.2% |
| economics | Gemini 3.5 Flash-Lite | 85.9% |
| economics | Qwen3.6 35B-A3B | 85.3% |
| economics | DeepSeek V4 Pro 0813 | 85.2% |
| economics | Qwen3.8 27B (FP8) | 85.1% |
| economics | DeepSeek V4 Pro | 84.1% |
| economics | Gemma 4 26B A4B | 82.6% |
| economics | Tencent HY3-Preview | 82.3% |
| economics | Claude Haiku 4.5 | 81.5% |
| economics | Gemini 2.5 Flash-Lite | 79.1% |
| economics | DeepSeek V4 Flash | 75.0% |
| engineering | Gemini 3.8 Flash | 86.8% |
| engineering | Gemini 3.7 Flash | 84.5% |
| engineering | GPT-5.5 | 83.3% |
| engineering | Gemini 3.5 Flash | 82.5% |
| engineering | Gemini 3 Flash Preview | 81.2% |
| engineering | Gemini 3.6 Flash | 80.5% |
| engineering | Gemini 3.1 Flash-Lite | 77.8% |
| engineering | Qwen3.6 35B-A3B | 77.5% |
| engineering | Qwen3.6 27B (BF16) | 77.3% |
| engineering | Gemma 4 31B | 77.1% |
| engineering | DeepSeek V4 Pro | 74.2% |
| engineering | Qwen3.8 27B (FP8) | 74.2% |
| engineering | Gemma 4 26B A4B | 73.5% |
| engineering | Qwen3.8 27B (BF16) | 73.2% |
| engineering | Gemini 3.5 Flash-Lite | 72.8% |
| engineering | Gemini 2.5 Flash | 71.6% |
| engineering | DeepSeek V4 Pro 0813 | 71.4% |
| engineering | Claude Haiku 4.5 | 64.7% |
| engineering | Tencent HY3-Preview | 64.4% |
| engineering | DeepSeek V4 Flash | 64.1% |
| engineering | Gemini 2.5 Flash-Lite | 58.4% |
| health | Gemini 3.8 Flash | 82.4% |
| health | Gemini 3.7 Flash | 82.1% |
| health | GPT-5.5 | 81.8% |
| health | Gemini 3 Flash Preview | 80.1% |
| health | Gemini 3.6 Flash | 79.3% |
| health | Gemini 3.5 Flash | 78.6% |
| health | DeepSeek V4 Pro 0813 | 75.1% |
| health | Qwen3.6 27B (BF16) | 75.1% |
| health | Gemini 3.1 Flash-Lite | 74.8% |
| health | DeepSeek V4 Pro | 74.2% |
| health | Gemma 4 31B | 74.1% |
| health | Gemini 3.5 Flash-Lite | 73.9% |
| health | Tencent HY3-Preview | 73.7% |
| health | Claude Haiku 4.5 | 72.9% |
| health | Qwen3.6 35B-A3B | 72.8% |
| health | Gemini 2.5 Flash | 72.5% |
| health | Qwen3.8 27B (BF16) | 71.8% |
| health | DeepSeek V4 Flash | 71.3% |
| health | Gemma 4 26B A4B | 71.2% |
| health | Qwen3.8 27B (FP8) | 70.9% |
| health | Gemini 2.5 Flash-Lite | 66.5% |
| history | Gemini 3 Flash Preview | 80.8% |
| history | Gemini 3.8 Flash | 80.8% |
| history | Gemini 3.5 Flash | 80.6% |
| history | Gemini 3.6 Flash | 80.1% |
| history | Gemini 3.7 Flash | 80.1% |
| history | GPT-5.5 | 77.7% |
| history | Gemini 3.5 Flash-Lite | 75.1% |
| history | Gemini 3.1 Flash-Lite | 74.5% |
| history | Gemma 4 31B | 74.5% |
| history | DeepSeek V4 Pro 0813 | 74.2% |
| history | Qwen3.6 27B (BF16) | 72.6% |
| history | DeepSeek V4 Pro | 71.9% |
| history | Tencent HY3-Preview | 71.4% |
| history | Gemini 2.5 Flash | 70.9% |
| history | Qwen3.8 27B (BF16) | 70.5% |
| history | Qwen3.8 27B (FP8) | 69.2% |
| history | Qwen3.6 35B-A3B | 69.0% |
| history | DeepSeek V4 Flash | 68.2% |
| history | Gemma 4 26B A4B | 66.7% |
| history | Claude Haiku 4.5 | 65.4% |
| history | Gemini 2.5 Flash-Lite | 59.3% |
| law | Gemini 3.8 Flash | 76.9% |
| law | Gemini 3.7 Flash | 76.6% |
| law | GPT-5.5 | 74.8% |
| law | Gemini 3 Flash Preview | 73.2% |
| law | Gemini 3.6 Flash | 73.2% |
| law | Gemini 3.5 Flash | 72.8% |
| law | Gemini 3.1 Flash-Lite | 62.8% |
| law | Gemma 4 31B | 59.0% |
| law | DeepSeek V4 Pro 0813 | 58.7% |
| law | Gemini 3.5 Flash-Lite | 58.2% |
| law | Qwen3.6 27B (BF16) | 57.1% |
| law | DeepSeek V4 Pro | 55.3% |
| law | Gemini 2.5 Flash | 54.0% |
| law | Qwen3.8 27B (FP8) | 53.8% |
| law | Qwen3.6 35B-A3B | 52.8% |
| law | Qwen3.8 27B (BF16) | 52.2% |
| law | Gemma 4 26B A4B | 50.7% |
| law | Claude Haiku 4.5 | 47.4% |
| law | DeepSeek V4 Flash | 46.1% |
| law | Tencent HY3-Preview | 45.8% |
| law | Gemini 2.5 Flash-Lite | 41.6% |
| math | Gemini 3.7 Flash | 95.4% |
| math | Gemini 3.8 Flash | 95.4% |
| math | Gemini 3.5 Flash | 94.6% |
| math | GPT-5.5 | 94.6% |
| math | Gemini 3 Flash Preview | 94.1% |
| math | Gemini 3.6 Flash | 93.9% |
| math | Qwen3.6 27B (BF16) | 93.7% |
| math | Gemma 4 31B | 93.5% |
| math | Qwen3.8 27B (FP8) | 93.4% |
| math | Qwen3.8 27B (BF16) | 93.3% |
| math | Qwen3.6 35B-A3B | 92.7% |
| math | Gemma 4 26B A4B | 92.0% |
| math | DeepSeek V4 Pro 0813 | 91.5% |
| math | DeepSeek V4 Pro | 90.9% |
| math | Gemini 2.5 Flash | 90.5% |
| math | Gemini 3.1 Flash-Lite | 90.2% |
| math | Gemini 3.5 Flash-Lite | 89.9% |
| math | DeepSeek V4 Flash | 88.7% |
| math | Claude Haiku 4.5 | 86.8% |
| math | Tencent HY3-Preview | 86.5% |
| math | Gemini 2.5 Flash-Lite | 82.8% |
| nothink biology | Gemini 3.1 Pro | 95.1% |
| nothink biology | Gemma 4 12B | 87.0% |
| nothink biology | Qwen3.5 9B | 85.6% |
| nothink biology | Qwen3 14B | 81.3% |
| nothink business | Gemini 3.1 Pro | 91.8% |
| nothink business | Gemma 4 12B | 79.8% |
| nothink business | Qwen3.5 9B | 78.1% |
| nothink business | Qwen3 14B | 70.3% |
| nothink chemistry | Gemini 3.1 Pro | 91.0% |
| nothink chemistry | Qwen3.5 9B | 85.0% |
| nothink chemistry | Gemma 4 12B | 79.9% |
| nothink chemistry | Qwen3 14B | 72.2% |
| nothink computer science | Gemini 3.1 Pro | 91.5% |
| nothink computer science | Gemma 4 12B | 79.3% |
| nothink computer science | Qwen3.5 9B | 78.0% |
| nothink computer science | Qwen3 14B | 71.5% |
| nothink economics | Gemini 3.1 Pro | 91.4% |
| nothink economics | Qwen3.5 9B | 78.9% |
| nothink economics | Gemma 4 12B | 78.7% |
| nothink economics | Qwen3 14B | 71.1% |
| nothink engineering | Gemini 3.1 Pro | 86.0% |
| nothink engineering | Qwen3.5 9B | 66.6% |
| nothink engineering | Gemma 4 12B | 65.1% |
| nothink engineering | Qwen3 14B | 55.9% |
| nothink health | Gemini 3.1 Pro | 82.8% |
| nothink health | Qwen3.5 9B | 65.9% |
| nothink health | Gemma 4 12B | 64.5% |
| nothink health | Qwen3 14B | 56.8% |
| nothink history | Gemini 3.1 Pro | 82.7% |
| nothink history | Qwen3.5 9B | 60.4% |
| nothink history | Gemma 4 12B | 57.2% |
| nothink history | Qwen3 14B | 48.8% |
| nothink law | Gemini 3.1 Pro | 78.6% |
| nothink law | Gemma 4 12B | 42.8% |
| nothink law | Qwen3.5 9B | 36.7% |
| nothink law | Qwen3 14B | 27.4% |
| nothink math | Gemini 3.1 Pro | 95.0% |
| nothink math | Qwen3.5 9B | 90.2% |
| nothink math | Gemma 4 12B | 90.1% |
| nothink math | Qwen3 14B | 82.2% |
| nothink other | Gemini 3.1 Pro | 85.1% |
| nothink other | Qwen3.5 9B | 61.6% |
| nothink other | Gemma 4 12B | 60.7% |
| nothink other | Qwen3 14B | 52.1% |
| nothink philosophy | Gemini 3.1 Pro | 86.0% |
| nothink philosophy | Gemma 4 12B | 60.3% |
| nothink philosophy | Qwen3.5 9B | 58.1% |
| nothink philosophy | Qwen3 14B | 49.3% |
| nothink physics | Gemini 3.1 Pro | 93.6% |
| nothink physics | Qwen3.5 9B | 84.7% |
| nothink physics | Gemma 4 12B | 81.9% |
| nothink physics | Qwen3 14B | 73.1% |
| nothink psychology | Gemini 3.1 Pro | 90.2% |
| nothink psychology | Gemma 4 12B | 76.3% |
| nothink psychology | Qwen3.5 9B | 74.2% |
| nothink psychology | Qwen3 14B | 66.0% |
| other | Gemini 3.8 Flash | 84.6% |
| other | GPT-5.5 | 84.3% |
| other | Gemini 3.7 Flash | 83.8% |
| other | Gemini 3 Flash Preview | 82.6% |
| other | Gemini 3.5 Flash | 82.4% |
| other | Gemini 3.6 Flash | 81.5% |
| other | DeepSeek V4 Pro | 77.3% |
| other | DeepSeek V4 Pro 0813 | 77.1% |
| other | Gemini 3.1 Flash-Lite | 76.7% |
| other | Gemini 3.5 Flash-Lite | 75.3% |
| other | Qwen3.6 27B (BF16) | 74.3% |
| other | Gemma 4 31B | 73.7% |
| other | Gemini 2.5 Flash | 73.5% |
| other | Tencent HY3-Preview | 72.1% |
| other | Qwen3.6 35B-A3B | 70.8% |
| other | Qwen3.8 27B (BF16) | 70.5% |
| other | Qwen3.8 27B (FP8) | 70.2% |
| other | Claude Haiku 4.5 | 69.5% |
| other | DeepSeek V4 Flash | 69.4% |
| other | Gemma 4 26B A4B | 67.5% |
| other | Gemini 2.5 Flash-Lite | 64.9% |
| philosophy | Gemini 3.8 Flash | 87.2% |
| philosophy | Gemini 3.7 Flash | 85.4% |
| philosophy | GPT-5.5 | 85.2% |
| philosophy | Gemini 3.5 Flash | 83.2% |
| philosophy | Gemini 3 Flash Preview | 82.4% |
| philosophy | Gemini 3.6 Flash | 82.4% |
| philosophy | DeepSeek V4 Pro 0813 | 75.7% |
| philosophy | Gemini 3.1 Flash-Lite | 75.2% |
| philosophy | Gemma 4 31B | 74.3% |
| philosophy | Qwen3.6 27B (BF16) | 73.9% |
| philosophy | DeepSeek V4 Pro | 73.7% |
| philosophy | Gemini 3.5 Flash-Lite | 72.9% |
| philosophy | DeepSeek V4 Flash | 71.9% |
| philosophy | Gemini 2.5 Flash | 70.9% |
| philosophy | Qwen3.6 35B-A3B | 69.5% |
| philosophy | Gemma 4 26B A4B | 69.1% |
| philosophy | Tencent HY3-Preview | 69.1% |
| philosophy | Qwen3.8 27B (FP8) | 68.5% |
| philosophy | Claude Haiku 4.5 | 67.9% |
| philosophy | Qwen3.8 27B (BF16) | 66.9% |
| philosophy | Gemini 2.5 Flash-Lite | 60.9% |
| physics | Gemini 3.8 Flash | 93.5% |
| physics | Gemini 3.7 Flash | 93.0% |
| physics | GPT-5.5 | 90.8% |
| physics | Gemini 3.5 Flash | 90.4% |
| physics | Gemini 3.6 Flash | 90.4% |
| physics | Gemini 3 Flash Preview | 90.3% |
| physics | Qwen3.6 27B (BF16) | 89.7% |
| physics | Gemma 4 31B | 89.0% |
| physics | Qwen3.8 27B (BF16) | 88.7% |
| physics | Qwen3.6 35B-A3B | 88.0% |
| physics | Qwen3.8 27B (FP8) | 87.7% |
| physics | Gemini 3.1 Flash-Lite | 87.6% |
| physics | DeepSeek V4 Pro | 86.8% |
| physics | Gemma 4 26B A4B | 86.1% |
| physics | Gemini 3.5 Flash-Lite | 85.9% |
| physics | Gemini 2.5 Flash | 85.5% |
| physics | DeepSeek V4 Pro 0813 | 85.1% |
| physics | DeepSeek V4 Flash | 84.7% |
| physics | Claude Haiku 4.5 | 81.3% |
| physics | Gemini 2.5 Flash-Lite | 76.6% |
| physics | Tencent HY3-Preview | 64.7% |
| psychology | Gemini 3.8 Flash | 89.2% |
| psychology | Gemini 3.5 Flash | 87.8% |
| psychology | Gemini 3 Flash Preview | 87.8% |
| psychology | Gemini 3.7 Flash | 87.5% |
| psychology | Gemini 3.6 Flash | 87.3% |
| psychology | GPT-5.5 | 86.3% |
| psychology | Gemini 3.5 Flash-Lite | 84.0% |
| psychology | DeepSeek V4 Pro 0813 | 83.9% |
| psychology | Gemma 4 31B | 83.6% |
| psychology | Gemini 3.1 Flash-Lite | 83.3% |
| psychology | DeepSeek V4 Pro | 83.3% |
| psychology | Gemini 2.5 Flash | 82.6% |
| psychology | Qwen3.6 27B (BF16) | 82.5% |
| psychology | Qwen3.8 27B (FP8) | 81.8% |
| psychology | DeepSeek V4 Flash | 80.8% |
| psychology | Tencent HY3-Preview | 80.5% |
| psychology | Qwen3.8 27B (BF16) | 79.5% |
| psychology | Claude Haiku 4.5 | 78.9% |
| psychology | Qwen3.6 35B-A3B | 78.7% |
| psychology | Gemma 4 26B A4B | 78.6% |
| psychology | Gemini 2.5 Flash-Lite | 75.1% |
Multi-step soft reasoning over long narrative contexts — murder mysteries, object placement and team allocation. Requires chaining clues across several paragraphs to reach the correct answer. Translated to German from the original English MuSR benchmark. Public results use the corrected v1.1 evaluation with 553 effective test cases.
| # | Model | Score | Latency | Tok in | Answer tok | Cost | Date |
|---|---|---|---|---|---|---|---|
| 1 | Claude Fable 5 🔒 reasoning | 89.5% | 20.28s | 3,008.8 | 1,338.2 | $54.71 | 2026-07-05 View result → |
| 2 | Gemini 3.8 Flash 🔒 reasoning | 89.0% | 20.21s | 1,484.9 | 462.7 | — | 2026-09-03 View result → |
| 3 | Gemini 3.7 Flash 🔒 reasoning | 88.6% | 8.23s | 1,484.9 | 372.5 | $2.37 | 2026-08-14 View result → |
| 4 | Gemini 3.1 Pro 🔒 reasoning | 88.1% | 12.14s | 1,464.9 | 510.8 | $10.49 | 2026-06-09 View result → |
| 5 | GPT-5.5 🔒 reasoning | 86.3% | 13.58s | 1,449.5 | 414.8 | $14.53 | 2026-06-13 View result → |
| 6 | Claude Opus 4.8 | 86.3% | 13.87s | 3,039.8 | 1,075.6 | $23.74 | 2026-06-13 View result → |
| 7 | Muse Glimmer 30B 🔒 reasoning | 86.1% | 10.15s | 1,500 | 511 | $1.00 | 2026-08-15 View result → |
| 8 | Qwen3.8 27B (BF16) | 86.1% | 40.69s | 1,489.1 | 1,772.5 | — | 2026-08-15 View result → |
| 9 | DeepSeek V4 Pro | 85.3% | 12.60s | 1,657.8 | 719.7 | $0.74 | 2026-06-23 View result → |
| 10 | DeepSeek V4 Flash 0423 | 85.3% | 16.41s | 1,659.8 | 801 | $0.16 | 2026-08-11 View result → |
| 11 | Tencent HY3 | 85.0% | 8.65s | 1,889.7 | 896.7 | — | 2026-08-30 View result → |
| 12 | Qwen3.6 27B (BF16) | 84.4% | 39.15s | 1,489.1 | 1,538.5 | — | 2026-08-15 View result → |
| 13 | Gemini 3.1 Flash-Lite | 84.4% | 2.95s | 1,467.3 | 568.5 | $0.73 | 2026-06-09 View result → |
| 14 | Gemini 3.5 Flash | 84.2% | 5.79s | 1,464.9 | 850.2 | $5.55 | 2026-06-09 View result → |
| 15 | Gemma 4 26B A4B | 83.7% | — | — | — | — | 2026-06-10 View result → |
| 16 | DeepSeek V4 Flash | 83.5% | 10.18s | 1,659.8 | 798.9 | $0.26 | 2026-06-10 View result → |
| 17 | Gemma 4 31B | 83.5% | 23.30s | 1,477.9 | 650.1 | $0.23 | 2026-06-11 View result → |
| 18 | Gemini 2.5 Flash | 83.3% | 6.28s | 1,464.9 | 1,077.4 | $1.77 | 2026-06-09 View result → |
| 19 | Gemini 3.5 Flash-Lite | 82.3% | 3.13s | 1,484.9 | 653 | $1.15 | 2026-08-30 View result → |
| 20 | DeepSeek V4 Pro 0813 | 82.3% | 9.50s | 1,682.4 | 580.3 | $0.67 | 2026-08-13 View result → |
| 21 | GLM-5.3 Flash 🔒 reasoning | 81.9% | 19.89s | 1,612.1 | 386.6 | $0.0038 | 2026-08-30 View result → |
| 22 | Gemma 4 12B | 81.6% | 195.50s | 1,477.9 | 707.4 | — | 2026-06-11 View result → |
| 23 | GLM-5.1 | 81.4% | — | — | — | — | 2026-06-10 View result → |
| 24 | MiMo V2.5 Pro | 80.9% | — | — | — | — | 2026-06-10 View result → |
| 25 | Qwen3.5 9B | 80.9% | 61.14s | 1,469.2 | 1,629.7 | $0.22 | 2026-06-11 View result → |
| 26 | Gemini 3.6 Flash | 80.3% | 4.01s | 1,464.9 | 663.7 | $4.05 | 2026-07-23 View result → |
| 27 | Tencent HY4-Preview | 80.1% | 16.08s | 1,730.6 | 929.8 | — | 2026-08-30 View result → |
| 28 | Gemini 3 Flash Preview | 79.3% | 5.41s | 1,464.9 | 622.9 | $1.47 | 2026-06-25 View result → |
| 29 | GPT-5.6 Terra | 78.7% | 7.62s | 1,445.7 | 382.9 | $2.26 | 2026-08-13 View result → |
| 30 | GPT-5.6 Luna | 75.8% | 3.75s | 1,445.7 | 322.2 | $0.21 | 2026-08-13 View result → |
| 31 | Qwen3 14B | 69.7% | 263.43s | 1,728.9 | 1,270.4 | — | 2026-06-11 View result → |
| Subject | Model | Score |
|---|---|---|
| murder mystery | Claude Fable 5 | 92.8% |
| murder mystery | Gemini 3.8 Flash | 91.6% |
| murder mystery | Gemini 3.7 Flash | 91.2% |
| murder mystery | GPT-5.5 | 90.0% |
| murder mystery | Gemini 3.1 Pro | 89.2% |
| murder mystery | Muse Glimmer 30B | 88.4% |
| murder mystery | Qwen3.8 27B (BF16) | 88.4% |
| murder mystery | Gemini 2.5 Flash | 87.6% |
| murder mystery | Claude Opus 4.8 | 87.6% |
| murder mystery | DeepSeek V4 Flash 0423 | 87.6% |
| murder mystery | Gemma 4 31B | 86.4% |
| murder mystery | DeepSeek V4 Pro | 86.4% |
| murder mystery | DeepSeek V4 Flash | 85.6% |
| murder mystery | Gemma 4 12B | 85.6% |
| murder mystery | DeepSeek V4 Pro 0813 | 85.6% |
| murder mystery | Gemma 4 26B A4B | 85.2% |
| murder mystery | Gemini 3.5 Flash-Lite | 85.2% |
| murder mystery | Gemini 3.1 Flash-Lite | 84.8% |
| murder mystery | Tencent HY3 | 84.8% |
| murder mystery | Gemini 3.5 Flash | 84.0% |
| murder mystery | Qwen3.6 27B (BF16) | 84.0% |
| murder mystery | GPT-5.6 Terra | 82.8% |
| murder mystery | GLM-5.3 Flash | 81.6% |
| murder mystery | GPT-5.6 Luna | 80.4% |
| murder mystery | GLM-5.1 | 80.0% |
| murder mystery | Gemini 3.6 Flash | 79.6% |
| murder mystery | MiMo V2.5 Pro | 78.8% |
| murder mystery | Gemini 3 Flash Preview | 78.8% |
| murder mystery | Tencent HY4-Preview | 78.0% |
| murder mystery | Qwen3.5 9B | 77.6% |
| murder mystery | Qwen3 14B | 76.8% |
| object placements | GPT-5.6 Luna | 82.8% |
| object placements | Gemini 3.5 Flash | 81.3% |
| object placements | GLM-5.1 | 81.3% |
| object placements | Claude Fable 5 | 81.3% |
| object placements | Gemini 3.1 Pro | 79.7% |
| object placements | Claude Opus 4.8 | 79.7% |
| object placements | Gemma 4 12B | 79.7% |
| object placements | Tencent HY4-Preview | 79.7% |
| object placements | MiMo V2.5 Pro | 76.6% |
| object placements | Gemma 4 31B | 76.6% |
| object placements | Gemini 3 Flash Preview | 76.6% |
| object placements | Muse Glimmer 30B | 76.6% |
| object placements | Tencent HY3 | 76.6% |
| object placements | Gemini 3.1 Flash-Lite | 75.0% |
| object placements | Gemini 2.5 Flash | 75.0% |
| object placements | DeepSeek V4 Flash | 75.0% |
| object placements | Gemini 3.5 Flash-Lite | 75.0% |
| object placements | Gemini 3.7 Flash | 75.0% |
| object placements | GPT-5.6 Terra | 75.0% |
| object placements | Qwen3.8 27B (BF16) | 75.0% |
| object placements | Gemini 3.8 Flash | 75.0% |
| object placements | Gemma 4 26B A4B | 73.4% |
| object placements | Qwen3.5 9B | 73.4% |
| object placements | Gemini 3.6 Flash | 73.4% |
| object placements | Qwen3 14B | 71.9% |
| object placements | DeepSeek V4 Flash 0423 | 71.9% |
| object placements | DeepSeek V4 Pro 0813 | 71.9% |
| object placements | Qwen3.6 27B (BF16) | 71.9% |
| object placements | DeepSeek V4 Pro | 70.3% |
| object placements | GPT-5.5 | 68.8% |
| object placements | GLM-5.3 Flash | 68.8% |
| team allocation | Gemini 3.8 Flash | 90.0% |
| team allocation | Gemini 3.7 Flash | 89.5% |
| team allocation | Gemini 3.1 Pro | 89.2% |
| team allocation | Claude Fable 5 | 88.4% |
| team allocation | Qwen3.6 27B (BF16) | 88.3% |
| team allocation | DeepSeek V4 Pro | 88.0% |
| team allocation | Tencent HY3 | 87.4% |
| team allocation | GPT-5.5 | 87.2% |
| team allocation | Claude Opus 4.8 | 86.8% |
| team allocation | Qwen3.8 27B (BF16) | 86.6% |
| team allocation | Gemini 3.1 Flash-Lite | 86.4% |
| team allocation | DeepSeek V4 Flash 0423 | 86.4% |
| team allocation | Muse Glimmer 30B | 86.2% |
| team allocation | Qwen3.5 9B | 86.0% |
| team allocation | GLM-5.3 Flash | 85.8% |
| team allocation | Gemini 3.5 Flash | 85.2% |
| team allocation | DeepSeek V4 Flash | 84.8% |
| team allocation | Gemma 4 26B A4B | 84.8% |
| team allocation | MiMo V2.5 Pro | 84.0% |
| team allocation | GLM-5.1 | 82.8% |
| team allocation | Gemini 3.6 Flash | 82.8% |
| team allocation | Tencent HY4-Preview | 82.4% |
| team allocation | Gemma 4 31B | 82.4% |
| team allocation | DeepSeek V4 Pro 0813 | 81.6% |
| team allocation | Gemini 2.5 Flash | 81.2% |
| team allocation | Gemini 3.5 Flash-Lite | 81.2% |
| team allocation | Gemini 3 Flash Preview | 80.4% |
| team allocation | Gemma 4 12B | 78.0% |
| team allocation | GPT-5.6 Terra | 75.3% |
| team allocation | GPT-5.6 Luna | 69.0% |
| team allocation | Qwen3 14B | 62.0% |
Native German social-media sentiment classification — positive, neutral or negative. Human-annotated German text, not translated. Run reasoning-off; headline score is MCC on the predicted label.
Native German linguistic acceptability — does the sentence read as grammatical German (ja / nein)? Built from clean vs. minimally-corrupted German sentences. Run reasoning-off.
TPS is decode speed — output tokens per second after the first token; higher is faster. TTFT is time to first token; lower is snappier. A snapshot, not a constant.
| # | Model | TPS | TTFT |
|---|---|---|---|
| 1 | gpt-oss-120b 🔒 | 1,641 | 0.31s |
| 2 | Gemini 3.1 Flash-Lite | 223 | 0.61s |
| 3 | Gemini 3.5 Flash | 181 | 0.75s |
| 4 | Gemini 3.5 Flash-Lite | 181 | 0.71s |
| 5 | Gemini 3.6 Flash | 166 | 0.64s |
| 6 | Qwen3.6 35B-A3B | 159 | 0.62s |
| 7 | Gemini 2.5 Flash | 159 | 0.41s |
| 8 | Gemini 3 Flash Preview | 147 | 0.98s |
| 9 | Claude Haiku 4.5 | 131 | 0.79s |
| 10 | DeepSeek V4 Flash | 115 | 0.92s |
| 11 | Tencent HY3-Preview | 107 | 2.69s |
| 12 | Gemma 4 26B A4B | 46 | 1.16s |
| 13 | GLM-5.1 | 32 | 0.65s |
| 14 | Qwen3 14B | 17 | 0.36s |
Average primary score across completed benchmarks against decode speed (output tokens per second). Up and to the right is better — smarter and faster. The gold line is the speed frontier: the best score available at each speed.
* scored on fewer benchmarks, so its average is not directly comparable to full-coverage models.
🔒 reasoning can't be disabled — its decode speed includes forced reasoning tokens, so it isn't directly comparable to the reasoning-off models.