German Language LLM Index
a PeerBench project
German-language LLM leaderboard

Which LLM speaks German best?

Independent German-language benchmarks for frontier and fast models — quality, latency, and cost, ranked for real deployment decisions.

Recommended picks

Start here if you need a decision.

Compare models
Best overall

Gemini 3.8 Flash

Highest German aggregate score

76.1 German score
Google 6/6 benchmarks $6.89/1k
Best open weights

GLM-5.3 Flash

Strongest open-weight model in the ranking

61.8 German score
Z.ai 5/6 benchmarks $0.001/1k
Fast strong model

Gemini 3.5 Flash

High German score with measured fast throughput

61.8 German score
Google 6/6 benchmarks $4.71/1k 181 tok/s
Best value

GLM-5.3 Flash

Most German score per listed token price

61.8 German score
Z.ai 5/6 benchmarks $0.001/1k
Scope
German-language only

No coding, vision, multimodal, or generic English arena score in the public aggregate.

Evidence
47 models · 6 benchmarks

Coverage is shown per model; partial averages are marked instead of hidden.

Benchmarks
Native + translated German

INCLUDE, GermEval, SB10K, ScaLA, MMLU-ProX, MMMLU, and MuSR.

Freshness
Updated 2026-09-04

Published rows are gated before they reach the static site.

Overall German ranking

Official German-language score from German benchmark evidence only. Partial coverage is shown in every row.

Coverage
Weight
Price
Speed
#
Model
Score
Cost
Speed
1
Google · Closed · $6.89/1k · —
76.1
$6.89
/1k questions
tok/s
2
Google · Closed · $2.44/1k · —
73.4
$2.44
/1k questions
tok/s
3
Google · Closed · $7.89/1k · —
70.2
$7.89
/1k questions
tok/s
4
Anthropic · Closed · $31.45/1k · —
69.2
$31.45
/1k questions
tok/s
5
Anthropic · Closed · $12.83/1k · —
63.3
$12.83
/1k questions
tok/s
6
GPT-5.5 6/6
OpenAI · Closed · $13.95/1k · —
62.5
$13.95
/1k questions
tok/s
7
Google · Closed · $4.71/1k · 181 tok/s
61.8
$4.71
/1k questions
181
tok/s
8
Z.ai · Open weights · $0.001/1k · —
61.8
$0.001
/1k questions
tok/s
9
Google · Closed · $3.73/1k · 166 tok/s
59.3
$3.73
/1k questions
166
tok/s
10
OpenAI · Closed · $1.39/1k · —
57.8
$1.39
/1k questions
tok/s
11
Google · Closed · $1.04/1k · 181 tok/s
56.8
$1.04
/1k questions
181
tok/s
12
Google · Closed · $1.45/1k · 147 tok/s
56.0
$1.45
/1k questions
147
tok/s
13
Alibaba · Open weights · —/1k · —
54.8
/1k questions
tok/s
14
Alibaba · Open weights · —/1k · —
54.4
/1k questions
tok/s
15
Tencent · Open weights · —/1k · —
54.2
/1k questions
tok/s
16
DeepSeek · Open weights · $0.356/1k · —
53.8
$0.356
/1k questions
tok/s
17
DeepSeek · Open weights · $0.075/1k · —
53.3
$0.075
/1k questions
tok/s
18
Google · Closed · $0.772/1k · 223 tok/s
52.9
$0.772
/1k questions
223
tok/s
19
Alibaba · Open weights · —/1k · —
52.7
/1k questions
tok/s
20
DeepSeek · Open weights · $0.109/1k · —
52.3
$0.109
/1k questions
tok/s
21
Alibaba · Closed · $1.17/1k · —
52.1
$1.17
/1k questions
tok/s
22
Tencent · Open weights · —/1k · —
51.3
/1k questions
tok/s
23
DeepSeek · Open weights · $0.662/1k · —
51.3
$0.662
/1k questions
tok/s
24
OpenAI · Closed · $0.135/1k · —
51.1
$0.135
/1k questions
tok/s
25
Google · Closed · $1.92/1k · 159 tok/s
50.7
$1.92
/1k questions
159
tok/s
26
Google · Open weights · $0.317/1k · —
49.5
$0.317
/1k questions
tok/s
27
DeepSeek · Open weights · $0.113/1k · 115 tok/s
49.0
$0.113
/1k questions
115
tok/s
28
Xiaomi · Open weights · $1.00/1k · —
47.6
$1.00
/1k questions
tok/s
29
GLM-5.1 5/6
Z.ai · Open weights · $1.13/1k · 32 tok/s
44.4
$1.13
/1k questions
32
tok/s
30
Google · Open weights · $0.318/1k · 46 tok/s
42.9
$0.318
/1k questions
46
tok/s
31
Google · Open weights · —/1k · —
42.7
/1k questions
tok/s
32
MiMo V2.5 3/6
Xiaomi · Open weights · $0.160/1k · —
41.6
$0.160
/1k questions
tok/s
33
Alibaba · Open weights · $0.879/1k · 159 tok/s
40.8
$0.879
/1k questions
159
tok/s
34
Tencent · Open weights · $0.191/1k · 107 tok/s
36.1
$0.191
/1k questions
107
tok/s
35
Alibaba · Open weights · $0.248/1k · —
34.1
$0.248
/1k questions
tok/s
36
Qwen3 14B 6/6
Alibaba · Open weights · —/1k · 17 tok/s
29.9
/1k questions
17
tok/s

Showing 36 models that ran at least 3 of 6 German-language benchmarks (11 excluded for thin coverage). Cost is effective benchmark cost per 1,000 questions; speed is output tokens/sec.

GermEval — German NER

Native German named-entity recognition — identify persons, locations, organisations and misc entities in German text, emitted as JSON. Scored with seqeval micro-F1 excluding the noisy MISC class. Run reasoning-off.

1,024 questions Named-entity recognition Native German GermEval (via EuroEval) ↗
# Model Score Latency Tok in Answer tok Cost Date
1 Claude Fable 5 🔒 reasoning
87.7%
5.52s 2,083.6 76.1 $30.07 2026-07-06 View result →
2 Gemini 3.1 Pro 🔒 reasoning
87.3%
4.86s 1,059.7 29.5 $10.04 2026-06-09 View result →
3 Gemini 3.8 Flash 🔒 reasoning
87.2%
2.78s 1,060 26.2 $3.19 2026-09-04 View result →
4 Gemini 3.7 Flash 🔒 reasoning
86.9%
1.10s 1,060 27.9 $1.01 2026-08-14 View result →
5 GLM-5.3 Flash 🔒 reasoning
86.4%
0.00s 1,143.8 27 2026-08-30 View result →
6 Gemini 3.6 Flash
86.2%
1.04s 1,060 38.2 $1.92 2026-07-23 View result →
7 Gemini 3.5 Flash
86.0%
0.81s 1,077 24.6 $1.88 2026-06-09 View result →
8 GPT-5.5 🔒 reasoning
85.9%
1.86s 1,140.7 29.1 $7.08 2026-06-13 View result →
9 Claude Opus 4.8
85.8%
2.12s 2,110.5 42 $11.88 2026-06-13 View result →
10 Gemini 3.5 Flash-Lite
85.4%
0.67s 1,060 46.3 $0.44 2026-08-30 View result →
11 Gemini 3 Flash Preview
84.7%
1.20s 1,060 24.9 $0.62 2026-06-25 View result →
12 GPT-5.6 Terra
83.3%
1.27s 1,121.7 27.6 $1.60 2026-08-12 View result →
13 Gemini 3.1 Flash-Lite
83.2%
0.52s 1,077 25.7 $0.63 2026-06-08 View result →
14 Qwen3.6 27B (BF16)
82.7%
2.79s 1,154.3 25.8 2026-08-15 View result →
15 Gemma 4 31B
82.6%
11.92s 1,153 27.8 $0.15 2026-06-11 View result →
16 DeepSeek V4 Pro
82.2%
1.74s 1,181.2 26.3 $2.02 2026-06-09 View result →
17 Tencent HY4-Preview
82.1%
1.86s 1,229.1 27.9 2026-08-30 View result →
18 Qwen3.7 Max
82.1%
1.64s 1,154.3 25.2 $1.57 2026-06-09 View result →
19 DeepSeek V4 Flash 0731
82.0%
0.86s 1,181.2 26.8 $0.10 2026-08-11 View result →
20 GPT-5.6 Luna
81.9%
0.95s 1,121.7 29.5 $0.16 2026-08-13 View result →
21 MiMo V2.5 Pro
81.9%
1.09s 1,533.5 26.4 $1.49 2026-06-09 View result →
22 Gemma 4 26B A4B
81.8%
1.14s 1,153 30.3 $0.19 2026-06-08 View result →
23 Gemini 2.5 Flash
81.8%
0.52s 1,077 24.5 $0.39 2026-06-09 View result →
24 DeepSeek V4 Pro 0813
81.7%
1.61s 1,181.2 25.9 $0.55 2026-08-13 View result →
25 Qwen3.8 27B (FP8)
81.6%
3.60s 1,186.3 26.4 2026-08-15 View result →
26 Qwen3.8 27B (BF16)
81.5%
2.81s 1,186.3 26.7 2026-08-15 View result →
27 MiMo V2.5
80.8%
0.86s 1,529.5 26.2 $0.20 2026-08-11 View result →
28 DeepSeek V4 Flash 0423
80.7%
0.89s 1,181.2 26.9 $0.11 2026-08-11 View result →
29 NVIDIA Nemotron 3 Ultra 550B-A55B
80.3%
1.38s 1,183.3 25.7 2026-08-12 View result →
30 DeepSeek V4 Flash
80.2%
0.77s 1,181.2 26.8 $0.18 2026-06-09 View result →
31 Qwen3.6 35B-A3B
79.9%
0.67s 1,154.3 27.3 $0.21 2026-06-09 View result →
32 Gemma 4 12B
79.6%
43.80s 1,153 28.1 2026-06-11 View result →
33 Tencent HY3-Preview
77.3%
2.50s 1,313.7 35.4 $0.09 2026-06-09 View result →
34 Tencent HY3
76.1%
1.68s 1,313.7 34.4 2026-08-30 View result →
35 Qwen3 14B
73.0%
21.12s 1,292.5 28.5 2026-06-11 View result →
36 Qwen3.5 9B
72.6%
11.77s 1,154.3 26.9 $0.12 2026-06-11 View result →
37 GLM-5.1
52.2%
2.20s 1,127.6 26.5 $1.71 2026-06-09 View result →

INCLUDE — German

Native German exam and licensing questions covering region-specific knowledge — history, law, civics and culture. Written by humans in German, not translated.

107 questions 4-option multiple choice Native German CohereLabs/include-base-44 ↗
# Model Score Latency Tok in Answer tok Cost Date
1 Gemini 3.7 Flash 🔒 reasoning
93.5%
4.23s 108.4 84.3 $0.12 2026-08-14 View result →
2 Gemini 3.8 Flash 🔒 reasoning
93.5%
5.11s 108.4 96.7 2026-09-03 View result →
3 GPT-5.6 Terra
88.8%
1.65s 110.9 54.3 $0.05 2026-08-13 View result →
4 DeepSeek V4 Pro 0813
87.9%
2.24s 121.2 93.8 $0.01 2026-08-13 View result →
5 Qwen3.8 27B (BF16)
87.9%
9.56s 120.7 490.5 2026-08-15 View result →
6 NVIDIA Nemotron 3 Ultra 550B-A55B
86.9%
3.91s 124.1 154.2 2026-08-15 View result →
7 Qwen3.8 27B (FP8)
86.9%
6.72s 120.7 451.1 2026-08-15 View result →
8 GLM-5.3 Flash 🔒 reasoning
86.9%
4.28s 124.5 81 $0.0003 2026-08-30 View result →
9 Tencent HY4-Preview
86.9%
3.99s 142.7 184.8 2026-08-30 View result →
10 GPT-5.6 Luna
86.0%
1.21s 110.9 48.7 $0.0043 2026-08-13 View result →
11 Muse Glimmer 30B 🔒 reasoning
86.0%
2.23s 158.7 61.1 $0.04 2026-08-15 View result →
12 Gemini 3.5 Flash-Lite
85.0%
1.40s 108.4 235.2 $0.07 2026-08-30 View result →
13 Qwen3.6 27B (BF16)
83.2%
9.19s 120.7 529.2 2026-08-15 View result →
14 Tencent HY3
83.2%
2.73s 147 150.3 2026-08-30 View result →
15 GLM-5.1
82.2%
2026-06-08 View result →
16 Gemini 3.1 Pro 🔒 reasoning
77.7%
4.24s 107.1 128.9 $0.72 2026-06-10 View result →
17 GPT-5.5 🔒 reasoning
74.8%
3.99s 128.2 50.2 $0.70 2026-06-12 View result →
18 Gemini 3.6 Flash
74.8%
1.37s 105.8 168.7 $0.19 2026-07-23 View result →
19 Gemini 3.5 Flash
74.1%
1.72s 107.1 205.4 $0.28 2026-06-10 View result →
20 MiniMax M2.7 🔒 reasoning
72.7%
356.6 Undated View result →
21 Claude Fable 5 🔒 reasoning
72.7%
7.72s 228.2 437 $3.33 2026-07-05 View result →
22 Claude Sonnet 5
72.7%
4.92s 229.9 263.5 $0.64 2026-07-21 View result →
23 Gemini 3.1 Flash-Lite
71.9%
0.89s 107.1 151 $0.04 2026-06-10 View result →
24 Claude Opus 4.8
71.9%
5.77s 256.2 343.7 $1.32 2026-06-12 View result →
25 Gemini 3 Flash Preview
71.9%
2.37s 106.4 170.5 $0.08 2026-06-25 View result →
26 Qwen3.7 Max
71.2%
2026-06-08 View result →
27 Gemini 2.5 Flash
70.5%
1.61s 105.8 193.6 $0.07 2026-06-02 View result →
28 DeepSeek V4 Pro
70.5%
3.44s 118.7 121.3 $0.08 2026-06-08 View result →
29 DeepSeek V4 Flash
70.5%
2.53s 118.626 111.288 $0.0066 2026-06-08 View result →
30 Kimi K2.6
69.8%
13.65s 138.7 527.1 $0.58 2026-06-08 View result →
31 Tencent HY3-Preview
69.1%
5.48s 143.9 222.4 $0.0086 2026-06-02 View result →
32 Gemma 4 12B
69.1%
27.64s 123.1 221.9 2026-06-11 View result →
33 Claude Haiku 4.5
68.3%
3.45s 154.6 347.4 $0.26 2026-06-02 View result →
34 Qwen3.6 35B-A3B
68.3%
2.72s 118.4 545.1 $0.08 2026-06-02 View result →
35 Gemma 4 31B
68.3%
4.56s 119.1 205.4 $0.01 2026-06-11 View result →
36 MiMo V2.5 Pro
67.6%
2026-06-08 View result →
37 gpt-oss-120b 🔒 reasoning
66.2%
1.71s 154.1 21.4 $0.0078 2026-05-29 View result →
38 Gemma 4 26B A4B
64.7%
3.71s 118.7 275.8 $0.03 2026-06-03 View result →
39 Qwen3.5 9B
64.7%
10.95s 122.4 378 $0.0096 2026-06-11 View result →
40 Grok 4.3
63.3%
0.60s 233.8 20.5 $0.04 2026-06-03 View result →
41 Qwen3 14B
63.3%
3.32s 135.4 46 2026-06-11 View result →
42 Ministral 14B
58.3%
0.41s 109.1 79.8 $0.01 2026-06-03 View result →
43 Gemma 3 12B
53.2%
5.00s 114.8 160.5 $0.0073 2026-06-03 View result →

MMLU-ProX — German

Hard academic questions across 14 subjects — STEM, law, health, economics, philosophy and more. Professionally translated to German, with up to ten answer options per question.

11,759 questions 10-option multiple choice Professional translation li-lab/MMLU-ProX ↗
# Model Score Latency Tok in Answer tok Cost Date
1 Gemini 3.1 Pro 🔒 reasoning
89.2%
3.46s 1,668.9 243 $95.16 2026-08-15 View result →
2 Gemini 3.8 Flash 🔒 reasoning
89.0%
7.24s 1,664.9 222.4 $104.01 2026-09-04 View result →
3 Gemini 3.7 Flash 🔒 reasoning
88.2%
4.60s 1,664.9 191.5 $34.62 2026-08-14 View result →
4 GPT-5.5 🔒 reasoning
87.2%
29.48s 1,594.7 145.5 $193.51 2026-07-01 View result →
5 Gemini 3.5 Flash
86.5%
2.23s 1,664.9 353.3 $66.76 2026-05-26 View result →
6 Gemini 3 Flash Preview
86.3%
3.26s 1,664.9 311.5 $20.78 2026-06-25 View result →
7 Gemini 3.6 Flash
86.0%
1.84s 1,664.9 259.2 $52.23 2026-07-23 View result →
8 Qwen3.6 27B (BF16)
82.3%
18.46s 1,689.5 932.8 2026-08-15 View result →
9 Gemini 3.1 Flash-Lite
82.2%
1.24s 1,666.1 304.5 $10.21 2026-06-03 View result →
10 Gemma 4 31B
82.1%
64.35s 1,702.9 502.5 $4.53 2026-06-11 View result →
11 DeepSeek V4 Pro
80.8%
3.00s 288.1 218.9 $3.71 2026-06-16 View result →
12 DeepSeek V4 Pro 0813
80.7%
2.92s 1,718.4 184.8 $3.44 2026-08-14 View result →
13 Gemini 3.5 Flash-Lite
80.2%
1.57s 1,664.9 307.2 $14.90 2026-07-23 View result →
14 Qwen3.6 35B-A3B
80.0%
4.67s 1,689.6 984 $13.36 2026-05-31 View result →
15 Qwen3.8 27B (FP8)
80.0%
19.03s 1,689.5 1,210.1 2026-08-15 View result →
16 Qwen3.8 27B (BF16)
79.9%
25.21s 1,689.5 1,198.3 2026-08-15 View result →
17 Gemini 2.5 Flash
79.6%
2.69s 1,664.9 786 $28.89 2026-05-26 View result →
18 Gemma 4 26B A4B
78.2%
7.77s 1,700.6 662.1 $4.64 2026-06-05 View result →
19 Claude Haiku 4.5
75.3%
3.76s 2,262 433.7 $52.10 2026-06-03 View result →
20 DeepSeek V4 Flash
74.9%
1.98s 1,718.7 162.2 $1.04 2026-06-16 View result →
21 Qwen3.5 9B
73.4%
40.41s 1,693.6 931.1 $3.63 2026-06-11 View result →
22 Tencent HY3-Preview
73.2%
6.14s 1,910.2 586.3 $2.76 2026-05-27 View result →
23 Gemma 4 12B
73.1%
92.81s 1,705.9 493.3 2026-06-11 View result →
24 Gemini 2.5 Flash-Lite
71.2%
2.48s 1,665.1 1,493.5 $8.95 2026-06-03 View result →
25 Qwen3 14B
64.1%
29.26s 1,887.3 413.7 2026-06-11 View result →
Show per-subject breakdown (350)
Subject Model Score
biology Gemini 3.7 Flash 94.0%
biology Gemini 3.8 Flash 93.7%
biology Gemini 3.5 Flash 92.9%
biology Gemini 3 Flash Preview 92.6%
biology Gemini 3.6 Flash 92.6%
biology GPT-5.5 92.3%
biology Gemma 4 31B 91.2%
biology DeepSeek V4 Pro 90.2%
biology Qwen3.8 27B (BF16) 90.2%
biology DeepSeek V4 Pro 0813 90.1%
biology Qwen3.8 27B (FP8) 89.9%
biology Qwen3.6 27B (BF16) 89.8%
biology Qwen3.6 35B-A3B 89.7%
biology Gemini 3.1 Flash-Lite 89.4%
biology Gemini 2.5 Flash 88.8%
biology DeepSeek V4 Flash 88.4%
biology Gemini 3.5 Flash-Lite 88.1%
biology Gemma 4 26B A4B 87.3%
biology Gemini 2.5 Flash-Lite 86.8%
biology Claude Haiku 4.5 85.9%
biology Tencent HY3-Preview 85.8%
business Gemini 3.8 Flash 92.3%
business GPT-5.5 91.1%
business Gemini 3.7 Flash 90.6%
business Gemini 3.5 Flash 89.6%
business Gemini 3 Flash Preview 89.5%
business Gemini 3.6 Flash 88.2%
business Qwen3.6 27B (BF16) 87.7%
business Gemma 4 31B 87.5%
business Qwen3.8 27B (FP8) 85.9%
business Gemini 3.1 Flash-Lite 85.3%
business DeepSeek V4 Pro 85.3%
business Qwen3.8 27B (BF16) 85.1%
business Gemini 3.5 Flash-Lite 84.4%
business Gemma 4 26B A4B 83.9%
business Gemini 2.5 Flash 83.5%
business Qwen3.6 35B-A3B 83.5%
business DeepSeek V4 Pro 0813 83.4%
business Claude Haiku 4.5 79.0%
business Tencent HY3-Preview 78.3%
business Gemini 2.5 Flash-Lite 77.7%
business DeepSeek V4 Flash 76.2%
chemistry Gemini 3.8 Flash 91.8%
chemistry Gemini 3.7 Flash 91.4%
chemistry Qwen3.6 27B (BF16) 90.2%
chemistry GPT-5.5 89.8%
chemistry Gemini 3 Flash Preview 89.2%
chemistry Gemini 3.5 Flash 89.1%
chemistry Gemini 3.6 Flash 88.3%
chemistry Qwen3.8 27B (BF16) 88.3%
chemistry Qwen3.6 35B-A3B 88.2%
chemistry Qwen3.8 27B (FP8) 87.7%
chemistry Gemini 3.1 Flash-Lite 87.1%
chemistry Gemini 2.5 Flash 86.9%
chemistry Gemma 4 31B 86.7%
chemistry DeepSeek V4 Pro 86.7%
chemistry DeepSeek V4 Pro 0813 85.9%
chemistry Gemma 4 26B A4B 85.4%
chemistry Gemini 3.5 Flash-Lite 83.8%
chemistry Claude Haiku 4.5 80.4%
chemistry Gemini 2.5 Flash-Lite 78.4%
chemistry Tencent HY3-Preview 75.6%
chemistry DeepSeek V4 Flash 73.3%
computer science Gemini 3.8 Flash 91.5%
computer science Gemini 3.7 Flash 90.7%
computer science GPT-5.5 90.0%
computer science Gemini 3.1 Flash-Lite 88.8%
computer science Gemini 3.6 Flash 88.5%
computer science Qwen3.6 27B (BF16) 88.5%
computer science Gemini 3.5 Flash 87.6%
computer science Gemini 3 Flash Preview 86.6%
computer science DeepSeek V4 Flash 85.9%
computer science Gemma 4 31B 85.9%
computer science DeepSeek V4 Pro 85.4%
computer science Gemini 2.5 Flash 85.1%
computer science Qwen3.6 35B-A3B 85.1%
computer science Qwen3.8 27B (BF16) 84.1%
computer science Gemma 4 26B A4B 83.9%
computer science DeepSeek V4 Pro 0813 83.6%
computer science Claude Haiku 4.5 83.4%
computer science Qwen3.8 27B (FP8) 83.4%
computer science Gemini 3.5 Flash-Lite 83.2%
computer science Gemini 2.5 Flash-Lite 77.8%
computer science Tencent HY3-Preview 76.8%
economics Gemini 3.8 Flash 91.1%
economics Gemini 3.7 Flash 89.9%
economics GPT-5.5 89.2%
economics Gemini 3.5 Flash 89.1%
economics Gemini 3 Flash Preview 88.7%
economics Gemini 3.6 Flash 88.7%
economics Gemini 3.1 Flash-Lite 87.3%
economics Gemma 4 31B 86.8%
economics Qwen3.8 27B (BF16) 86.5%
economics Gemini 2.5 Flash 86.4%
economics Qwen3.6 27B (BF16) 86.2%
economics Gemini 3.5 Flash-Lite 85.9%
economics Qwen3.6 35B-A3B 85.3%
economics DeepSeek V4 Pro 0813 85.2%
economics Qwen3.8 27B (FP8) 85.1%
economics DeepSeek V4 Pro 84.1%
economics Gemma 4 26B A4B 82.6%
economics Tencent HY3-Preview 82.3%
economics Claude Haiku 4.5 81.5%
economics Gemini 2.5 Flash-Lite 79.1%
economics DeepSeek V4 Flash 75.0%
engineering Gemini 3.8 Flash 86.8%
engineering Gemini 3.7 Flash 84.5%
engineering GPT-5.5 83.3%
engineering Gemini 3.5 Flash 82.5%
engineering Gemini 3 Flash Preview 81.2%
engineering Gemini 3.6 Flash 80.5%
engineering Gemini 3.1 Flash-Lite 77.8%
engineering Qwen3.6 35B-A3B 77.5%
engineering Qwen3.6 27B (BF16) 77.3%
engineering Gemma 4 31B 77.1%
engineering DeepSeek V4 Pro 74.2%
engineering Qwen3.8 27B (FP8) 74.2%
engineering Gemma 4 26B A4B 73.5%
engineering Qwen3.8 27B (BF16) 73.2%
engineering Gemini 3.5 Flash-Lite 72.8%
engineering Gemini 2.5 Flash 71.6%
engineering DeepSeek V4 Pro 0813 71.4%
engineering Claude Haiku 4.5 64.7%
engineering Tencent HY3-Preview 64.4%
engineering DeepSeek V4 Flash 64.1%
engineering Gemini 2.5 Flash-Lite 58.4%
health Gemini 3.8 Flash 82.4%
health Gemini 3.7 Flash 82.1%
health GPT-5.5 81.8%
health Gemini 3 Flash Preview 80.1%
health Gemini 3.6 Flash 79.3%
health Gemini 3.5 Flash 78.6%
health DeepSeek V4 Pro 0813 75.1%
health Qwen3.6 27B (BF16) 75.1%
health Gemini 3.1 Flash-Lite 74.8%
health DeepSeek V4 Pro 74.2%
health Gemma 4 31B 74.1%
health Gemini 3.5 Flash-Lite 73.9%
health Tencent HY3-Preview 73.7%
health Claude Haiku 4.5 72.9%
health Qwen3.6 35B-A3B 72.8%
health Gemini 2.5 Flash 72.5%
health Qwen3.8 27B (BF16) 71.8%
health DeepSeek V4 Flash 71.3%
health Gemma 4 26B A4B 71.2%
health Qwen3.8 27B (FP8) 70.9%
health Gemini 2.5 Flash-Lite 66.5%
history Gemini 3 Flash Preview 80.8%
history Gemini 3.8 Flash 80.8%
history Gemini 3.5 Flash 80.6%
history Gemini 3.6 Flash 80.1%
history Gemini 3.7 Flash 80.1%
history GPT-5.5 77.7%
history Gemini 3.5 Flash-Lite 75.1%
history Gemini 3.1 Flash-Lite 74.5%
history Gemma 4 31B 74.5%
history DeepSeek V4 Pro 0813 74.2%
history Qwen3.6 27B (BF16) 72.6%
history DeepSeek V4 Pro 71.9%
history Tencent HY3-Preview 71.4%
history Gemini 2.5 Flash 70.9%
history Qwen3.8 27B (BF16) 70.5%
history Qwen3.8 27B (FP8) 69.2%
history Qwen3.6 35B-A3B 69.0%
history DeepSeek V4 Flash 68.2%
history Gemma 4 26B A4B 66.7%
history Claude Haiku 4.5 65.4%
history Gemini 2.5 Flash-Lite 59.3%
law Gemini 3.8 Flash 76.9%
law Gemini 3.7 Flash 76.6%
law GPT-5.5 74.8%
law Gemini 3 Flash Preview 73.2%
law Gemini 3.6 Flash 73.2%
law Gemini 3.5 Flash 72.8%
law Gemini 3.1 Flash-Lite 62.8%
law Gemma 4 31B 59.0%
law DeepSeek V4 Pro 0813 58.7%
law Gemini 3.5 Flash-Lite 58.2%
law Qwen3.6 27B (BF16) 57.1%
law DeepSeek V4 Pro 55.3%
law Gemini 2.5 Flash 54.0%
law Qwen3.8 27B (FP8) 53.8%
law Qwen3.6 35B-A3B 52.8%
law Qwen3.8 27B (BF16) 52.2%
law Gemma 4 26B A4B 50.7%
law Claude Haiku 4.5 47.4%
law DeepSeek V4 Flash 46.1%
law Tencent HY3-Preview 45.8%
law Gemini 2.5 Flash-Lite 41.6%
math Gemini 3.7 Flash 95.4%
math Gemini 3.8 Flash 95.4%
math Gemini 3.5 Flash 94.6%
math GPT-5.5 94.6%
math Gemini 3 Flash Preview 94.1%
math Gemini 3.6 Flash 93.9%
math Qwen3.6 27B (BF16) 93.7%
math Gemma 4 31B 93.5%
math Qwen3.8 27B (FP8) 93.4%
math Qwen3.8 27B (BF16) 93.3%
math Qwen3.6 35B-A3B 92.7%
math Gemma 4 26B A4B 92.0%
math DeepSeek V4 Pro 0813 91.5%
math DeepSeek V4 Pro 90.9%
math Gemini 2.5 Flash 90.5%
math Gemini 3.1 Flash-Lite 90.2%
math Gemini 3.5 Flash-Lite 89.9%
math DeepSeek V4 Flash 88.7%
math Claude Haiku 4.5 86.8%
math Tencent HY3-Preview 86.5%
math Gemini 2.5 Flash-Lite 82.8%
nothink biology Gemini 3.1 Pro 95.1%
nothink biology Gemma 4 12B 87.0%
nothink biology Qwen3.5 9B 85.6%
nothink biology Qwen3 14B 81.3%
nothink business Gemini 3.1 Pro 91.8%
nothink business Gemma 4 12B 79.8%
nothink business Qwen3.5 9B 78.1%
nothink business Qwen3 14B 70.3%
nothink chemistry Gemini 3.1 Pro 91.0%
nothink chemistry Qwen3.5 9B 85.0%
nothink chemistry Gemma 4 12B 79.9%
nothink chemistry Qwen3 14B 72.2%
nothink computer science Gemini 3.1 Pro 91.5%
nothink computer science Gemma 4 12B 79.3%
nothink computer science Qwen3.5 9B 78.0%
nothink computer science Qwen3 14B 71.5%
nothink economics Gemini 3.1 Pro 91.4%
nothink economics Qwen3.5 9B 78.9%
nothink economics Gemma 4 12B 78.7%
nothink economics Qwen3 14B 71.1%
nothink engineering Gemini 3.1 Pro 86.0%
nothink engineering Qwen3.5 9B 66.6%
nothink engineering Gemma 4 12B 65.1%
nothink engineering Qwen3 14B 55.9%
nothink health Gemini 3.1 Pro 82.8%
nothink health Qwen3.5 9B 65.9%
nothink health Gemma 4 12B 64.5%
nothink health Qwen3 14B 56.8%
nothink history Gemini 3.1 Pro 82.7%
nothink history Qwen3.5 9B 60.4%
nothink history Gemma 4 12B 57.2%
nothink history Qwen3 14B 48.8%
nothink law Gemini 3.1 Pro 78.6%
nothink law Gemma 4 12B 42.8%
nothink law Qwen3.5 9B 36.7%
nothink law Qwen3 14B 27.4%
nothink math Gemini 3.1 Pro 95.0%
nothink math Qwen3.5 9B 90.2%
nothink math Gemma 4 12B 90.1%
nothink math Qwen3 14B 82.2%
nothink other Gemini 3.1 Pro 85.1%
nothink other Qwen3.5 9B 61.6%
nothink other Gemma 4 12B 60.7%
nothink other Qwen3 14B 52.1%
nothink philosophy Gemini 3.1 Pro 86.0%
nothink philosophy Gemma 4 12B 60.3%
nothink philosophy Qwen3.5 9B 58.1%
nothink philosophy Qwen3 14B 49.3%
nothink physics Gemini 3.1 Pro 93.6%
nothink physics Qwen3.5 9B 84.7%
nothink physics Gemma 4 12B 81.9%
nothink physics Qwen3 14B 73.1%
nothink psychology Gemini 3.1 Pro 90.2%
nothink psychology Gemma 4 12B 76.3%
nothink psychology Qwen3.5 9B 74.2%
nothink psychology Qwen3 14B 66.0%
other Gemini 3.8 Flash 84.6%
other GPT-5.5 84.3%
other Gemini 3.7 Flash 83.8%
other Gemini 3 Flash Preview 82.6%
other Gemini 3.5 Flash 82.4%
other Gemini 3.6 Flash 81.5%
other DeepSeek V4 Pro 77.3%
other DeepSeek V4 Pro 0813 77.1%
other Gemini 3.1 Flash-Lite 76.7%
other Gemini 3.5 Flash-Lite 75.3%
other Qwen3.6 27B (BF16) 74.3%
other Gemma 4 31B 73.7%
other Gemini 2.5 Flash 73.5%
other Tencent HY3-Preview 72.1%
other Qwen3.6 35B-A3B 70.8%
other Qwen3.8 27B (BF16) 70.5%
other Qwen3.8 27B (FP8) 70.2%
other Claude Haiku 4.5 69.5%
other DeepSeek V4 Flash 69.4%
other Gemma 4 26B A4B 67.5%
other Gemini 2.5 Flash-Lite 64.9%
philosophy Gemini 3.8 Flash 87.2%
philosophy Gemini 3.7 Flash 85.4%
philosophy GPT-5.5 85.2%
philosophy Gemini 3.5 Flash 83.2%
philosophy Gemini 3 Flash Preview 82.4%
philosophy Gemini 3.6 Flash 82.4%
philosophy DeepSeek V4 Pro 0813 75.7%
philosophy Gemini 3.1 Flash-Lite 75.2%
philosophy Gemma 4 31B 74.3%
philosophy Qwen3.6 27B (BF16) 73.9%
philosophy DeepSeek V4 Pro 73.7%
philosophy Gemini 3.5 Flash-Lite 72.9%
philosophy DeepSeek V4 Flash 71.9%
philosophy Gemini 2.5 Flash 70.9%
philosophy Qwen3.6 35B-A3B 69.5%
philosophy Gemma 4 26B A4B 69.1%
philosophy Tencent HY3-Preview 69.1%
philosophy Qwen3.8 27B (FP8) 68.5%
philosophy Claude Haiku 4.5 67.9%
philosophy Qwen3.8 27B (BF16) 66.9%
philosophy Gemini 2.5 Flash-Lite 60.9%
physics Gemini 3.8 Flash 93.5%
physics Gemini 3.7 Flash 93.0%
physics GPT-5.5 90.8%
physics Gemini 3.5 Flash 90.4%
physics Gemini 3.6 Flash 90.4%
physics Gemini 3 Flash Preview 90.3%
physics Qwen3.6 27B (BF16) 89.7%
physics Gemma 4 31B 89.0%
physics Qwen3.8 27B (BF16) 88.7%
physics Qwen3.6 35B-A3B 88.0%
physics Qwen3.8 27B (FP8) 87.7%
physics Gemini 3.1 Flash-Lite 87.6%
physics DeepSeek V4 Pro 86.8%
physics Gemma 4 26B A4B 86.1%
physics Gemini 3.5 Flash-Lite 85.9%
physics Gemini 2.5 Flash 85.5%
physics DeepSeek V4 Pro 0813 85.1%
physics DeepSeek V4 Flash 84.7%
physics Claude Haiku 4.5 81.3%
physics Gemini 2.5 Flash-Lite 76.6%
physics Tencent HY3-Preview 64.7%
psychology Gemini 3.8 Flash 89.2%
psychology Gemini 3.5 Flash 87.8%
psychology Gemini 3 Flash Preview 87.8%
psychology Gemini 3.7 Flash 87.5%
psychology Gemini 3.6 Flash 87.3%
psychology GPT-5.5 86.3%
psychology Gemini 3.5 Flash-Lite 84.0%
psychology DeepSeek V4 Pro 0813 83.9%
psychology Gemma 4 31B 83.6%
psychology Gemini 3.1 Flash-Lite 83.3%
psychology DeepSeek V4 Pro 83.3%
psychology Gemini 2.5 Flash 82.6%
psychology Qwen3.6 27B (BF16) 82.5%
psychology Qwen3.8 27B (FP8) 81.8%
psychology DeepSeek V4 Flash 80.8%
psychology Tencent HY3-Preview 80.5%
psychology Qwen3.8 27B (BF16) 79.5%
psychology Claude Haiku 4.5 78.9%
psychology Qwen3.6 35B-A3B 78.7%
psychology Gemma 4 26B A4B 78.6%
psychology Gemini 2.5 Flash-Lite 75.1%

MuSR — German

Multi-step soft reasoning over long narrative contexts — murder mysteries, object placement and team allocation. Requires chaining clues across several paragraphs to reach the correct answer. Translated to German from the original English MuSR benchmark. Public results use the corrected v1.1 evaluation with 553 effective test cases.

553 questions 2–5 option multiple choice German translation TAUR-Lab/MuSR ↗
# Model Score Latency Tok in Answer tok Cost Date
1 Claude Fable 5 🔒 reasoning
89.5%
20.28s 3,008.8 1,338.2 $54.71 2026-07-05 View result →
2 Gemini 3.8 Flash 🔒 reasoning
89.0%
20.21s 1,484.9 462.7 2026-09-03 View result →
3 Gemini 3.7 Flash 🔒 reasoning
88.6%
8.23s 1,484.9 372.5 $2.37 2026-08-14 View result →
4 Gemini 3.1 Pro 🔒 reasoning
88.1%
12.14s 1,464.9 510.8 $10.49 2026-06-09 View result →
5 GPT-5.5 🔒 reasoning
86.3%
13.58s 1,449.5 414.8 $14.53 2026-06-13 View result →
6 Claude Opus 4.8
86.3%
13.87s 3,039.8 1,075.6 $23.74 2026-06-13 View result →
7 Muse Glimmer 30B 🔒 reasoning
86.1%
10.15s 1,500 511 $1.00 2026-08-15 View result →
8 Qwen3.8 27B (BF16)
86.1%
40.69s 1,489.1 1,772.5 2026-08-15 View result →
9 DeepSeek V4 Pro
85.3%
12.60s 1,657.8 719.7 $0.74 2026-06-23 View result →
10 DeepSeek V4 Flash 0423
85.3%
16.41s 1,659.8 801 $0.16 2026-08-11 View result →
11 Tencent HY3
85.0%
8.65s 1,889.7 896.7 2026-08-30 View result →
12 Qwen3.6 27B (BF16)
84.4%
39.15s 1,489.1 1,538.5 2026-08-15 View result →
13 Gemini 3.1 Flash-Lite
84.4%
2.95s 1,467.3 568.5 $0.73 2026-06-09 View result →
14 Gemini 3.5 Flash
84.2%
5.79s 1,464.9 850.2 $5.55 2026-06-09 View result →
15 Gemma 4 26B A4B
83.7%
2026-06-10 View result →
16 DeepSeek V4 Flash
83.5%
10.18s 1,659.8 798.9 $0.26 2026-06-10 View result →
17 Gemma 4 31B
83.5%
23.30s 1,477.9 650.1 $0.23 2026-06-11 View result →
18 Gemini 2.5 Flash
83.3%
6.28s 1,464.9 1,077.4 $1.77 2026-06-09 View result →
19 Gemini 3.5 Flash-Lite
82.3%
3.13s 1,484.9 653 $1.15 2026-08-30 View result →
20 DeepSeek V4 Pro 0813
82.3%
9.50s 1,682.4 580.3 $0.67 2026-08-13 View result →
21 GLM-5.3 Flash 🔒 reasoning
81.9%
19.89s 1,612.1 386.6 $0.0038 2026-08-30 View result →
22 Gemma 4 12B
81.6%
195.50s 1,477.9 707.4 2026-06-11 View result →
23 GLM-5.1
81.4%
2026-06-10 View result →
24 MiMo V2.5 Pro
80.9%
2026-06-10 View result →
25 Qwen3.5 9B
80.9%
61.14s 1,469.2 1,629.7 $0.22 2026-06-11 View result →
26 Gemini 3.6 Flash
80.3%
4.01s 1,464.9 663.7 $4.05 2026-07-23 View result →
27 Tencent HY4-Preview
80.1%
16.08s 1,730.6 929.8 2026-08-30 View result →
28 Gemini 3 Flash Preview
79.3%
5.41s 1,464.9 622.9 $1.47 2026-06-25 View result →
29 GPT-5.6 Terra
78.7%
7.62s 1,445.7 382.9 $2.26 2026-08-13 View result →
30 GPT-5.6 Luna
75.8%
3.75s 1,445.7 322.2 $0.21 2026-08-13 View result →
31 Qwen3 14B
69.7%
263.43s 1,728.9 1,270.4 2026-06-11 View result →
Show per-subject breakdown (93)
Subject Model Score
murder mystery Claude Fable 5 92.8%
murder mystery Gemini 3.8 Flash 91.6%
murder mystery Gemini 3.7 Flash 91.2%
murder mystery GPT-5.5 90.0%
murder mystery Gemini 3.1 Pro 89.2%
murder mystery Muse Glimmer 30B 88.4%
murder mystery Qwen3.8 27B (BF16) 88.4%
murder mystery Gemini 2.5 Flash 87.6%
murder mystery Claude Opus 4.8 87.6%
murder mystery DeepSeek V4 Flash 0423 87.6%
murder mystery Gemma 4 31B 86.4%
murder mystery DeepSeek V4 Pro 86.4%
murder mystery DeepSeek V4 Flash 85.6%
murder mystery Gemma 4 12B 85.6%
murder mystery DeepSeek V4 Pro 0813 85.6%
murder mystery Gemma 4 26B A4B 85.2%
murder mystery Gemini 3.5 Flash-Lite 85.2%
murder mystery Gemini 3.1 Flash-Lite 84.8%
murder mystery Tencent HY3 84.8%
murder mystery Gemini 3.5 Flash 84.0%
murder mystery Qwen3.6 27B (BF16) 84.0%
murder mystery GPT-5.6 Terra 82.8%
murder mystery GLM-5.3 Flash 81.6%
murder mystery GPT-5.6 Luna 80.4%
murder mystery GLM-5.1 80.0%
murder mystery Gemini 3.6 Flash 79.6%
murder mystery MiMo V2.5 Pro 78.8%
murder mystery Gemini 3 Flash Preview 78.8%
murder mystery Tencent HY4-Preview 78.0%
murder mystery Qwen3.5 9B 77.6%
murder mystery Qwen3 14B 76.8%
object placements GPT-5.6 Luna 82.8%
object placements Gemini 3.5 Flash 81.3%
object placements GLM-5.1 81.3%
object placements Claude Fable 5 81.3%
object placements Gemini 3.1 Pro 79.7%
object placements Claude Opus 4.8 79.7%
object placements Gemma 4 12B 79.7%
object placements Tencent HY4-Preview 79.7%
object placements MiMo V2.5 Pro 76.6%
object placements Gemma 4 31B 76.6%
object placements Gemini 3 Flash Preview 76.6%
object placements Muse Glimmer 30B 76.6%
object placements Tencent HY3 76.6%
object placements Gemini 3.1 Flash-Lite 75.0%
object placements Gemini 2.5 Flash 75.0%
object placements DeepSeek V4 Flash 75.0%
object placements Gemini 3.5 Flash-Lite 75.0%
object placements Gemini 3.7 Flash 75.0%
object placements GPT-5.6 Terra 75.0%
object placements Qwen3.8 27B (BF16) 75.0%
object placements Gemini 3.8 Flash 75.0%
object placements Gemma 4 26B A4B 73.4%
object placements Qwen3.5 9B 73.4%
object placements Gemini 3.6 Flash 73.4%
object placements Qwen3 14B 71.9%
object placements DeepSeek V4 Flash 0423 71.9%
object placements DeepSeek V4 Pro 0813 71.9%
object placements Qwen3.6 27B (BF16) 71.9%
object placements DeepSeek V4 Pro 70.3%
object placements GPT-5.5 68.8%
object placements GLM-5.3 Flash 68.8%
team allocation Gemini 3.8 Flash 90.0%
team allocation Gemini 3.7 Flash 89.5%
team allocation Gemini 3.1 Pro 89.2%
team allocation Claude Fable 5 88.4%
team allocation Qwen3.6 27B (BF16) 88.3%
team allocation DeepSeek V4 Pro 88.0%
team allocation Tencent HY3 87.4%
team allocation GPT-5.5 87.2%
team allocation Claude Opus 4.8 86.8%
team allocation Qwen3.8 27B (BF16) 86.6%
team allocation Gemini 3.1 Flash-Lite 86.4%
team allocation DeepSeek V4 Flash 0423 86.4%
team allocation Muse Glimmer 30B 86.2%
team allocation Qwen3.5 9B 86.0%
team allocation GLM-5.3 Flash 85.8%
team allocation Gemini 3.5 Flash 85.2%
team allocation DeepSeek V4 Flash 84.8%
team allocation Gemma 4 26B A4B 84.8%
team allocation MiMo V2.5 Pro 84.0%
team allocation GLM-5.1 82.8%
team allocation Gemini 3.6 Flash 82.8%
team allocation Tencent HY4-Preview 82.4%
team allocation Gemma 4 31B 82.4%
team allocation DeepSeek V4 Pro 0813 81.6%
team allocation Gemini 2.5 Flash 81.2%
team allocation Gemini 3.5 Flash-Lite 81.2%
team allocation Gemini 3 Flash Preview 80.4%
team allocation Gemma 4 12B 78.0%
team allocation GPT-5.6 Terra 75.3%
team allocation GPT-5.6 Luna 69.0%
team allocation Qwen3 14B 62.0%

SB10K — German sentiment

Native German social-media sentiment classification — positive, neutral or negative. Human-annotated German text, not translated. Run reasoning-off; headline score is MCC on the predicted label.

1,024 questions 3-class sentiment Native German SB10K (via EuroEval) ↗
# Model Score Latency Tok in Answer tok Cost Date
1 Gemini 3.1 Pro 🔒 reasoning
70.1%
3.30s 625.7 1.7 $3.81 2026-06-09 View result →
2 Gemini 3.7 Flash 🔒 reasoning
67.5%
0.91s 625.7 1.7 $0.53 2026-08-14 View result →
3 Claude Fable 5 🔒 reasoning
67.3%
4.60s 1,352.6 19.3 $18.05 2026-07-06 View result →
4 Gemini 3.8 Flash 🔒 reasoning
66.9%
1.64s 625.7 1.7 $1.88 2026-09-04 View result →
5 Claude Opus 4.8
64.5%
1.88s 1,372.7 5 $7.16 2026-06-13 View result →
6 Gemini 3.6 Flash
64.5%
0.93s 625.7 1.7 $0.97 2026-07-23 View result →
7 Gemini 3.5 Flash
64.3%
0.67s 650.6 1.7 $1.02 2026-06-09 View result →
8 Gemini 3 Flash Preview
63.6%
1.04s 625.7 1.7 $0.33 2026-06-25 View result →
9 GLM-5.3 Flash 🔒 reasoning
63.6%
3.42s 785.3 4 2026-08-30 View result →
10 Qwen3.7 Max
63.2%
0.99s 771.9 1.7 $0.99 2026-06-09 View result →
11 GPT-5.6 Terra
62.7%
1.01s 741.3 5.8 $0.79 2026-08-12 View result →
12 Qwen3.8 27B (FP8)
62.7%
2.27s 819.9 2.7 2026-08-15 View result →
13 Qwen3.8 27B (BF16)
62.6%
1.62s 819.9 2.7 2026-08-15 View result →
14 Gemini 3.5 Flash-Lite
62.5%
0.56s 625.7 1.7 $0.20 2026-08-30 View result →
15 GPT-5.5 🔒 reasoning
62.5%
1.36s 756.9 5.9 $4.12 2026-06-13 View result →
16 Tencent HY4-Preview
62.4%
1.57s 927.5 2.6 2026-08-30 View result →
17 Tencent HY3
62.4%
1.43s 888.4 2.6 2026-08-30 View result →
18 DeepSeek V4 Flash
62.4%
0.54s 778.6 2.7 $0.11 2026-06-09 View result →
19 Gemma 4 12B
62.2%
21.45s 758.7 2.7 2026-06-11 View result →
20 DeepSeek V4 Flash 0423
61.9%
0.64s 778.6 2.7 $0.07 2026-08-11 View result →
21 GPT-5.6 Luna
61.9%
0.84s 741.3 5.7 $0.08 2026-08-13 View result →
22 MiMo V2.5 Pro
61.7%
0.45s 1,140.3 2.8 $0.99 2026-06-09 View result →
23 Gemini 2.5 Flash
61.6%
0.40s 650.7 1.8 $0.20 2026-06-09 View result →
24 DeepSeek V4 Flash 0731
61.2%
0.67s 778.6 2.7 $0.06 2026-08-11 View result →
25 MiMo V2.5
61.2%
0.62s 1,136.3 2.8 $0.14 2026-08-11 View result →
26 Gemma 4 31B
61.0%
5.46s 758.7 2.8 $0.09 2026-06-11 View result →
27 Qwen3.6 27B (BF16)
61.0%
1.59s 771.9 2.7 2026-08-15 View result →
28 DeepSeek V4 Pro 0813
60.8%
1.35s 778.6 1.7 $0.35 2026-08-13 View result →
29 Gemini 3.1 Flash-Lite
60.4%
0.46s 650.7 1.8 $0.34 2026-06-08 View result →
30 DeepSeek V4 Pro
60.0%
1.42s 778.6 2.1 $1.28 2026-06-09 View result →
31 Qwen3.6 35B-A3B
58.4%
0.45s 771.9 2.7 $0.12 2026-06-09 View result →
32 Qwen3 14B
56.1%
13.26s 899.3 2.7 2026-06-11 View result →
33 Qwen3.5 9B
55.7%
6.96s 771.9 2.6 $0.08 2026-06-11 View result →
34 Tencent HY3-Preview
38.9%
2.35s 886.7 2.7 $0.06 2026-06-09 View result →
35 GLM-5.1
23.1%
1.88s 765.6 2.6 $1.09 2026-06-09 View result →
36 Gemma 4 26B A4B
15.2%
0.70s 758.8 2.8 $0.12 2026-06-08 View result →

ScaLA — German acceptability

Native German linguistic acceptability — does the sentence read as grammatical German (ja / nein)? Built from clean vs. minimally-corrupted German sentences. Run reasoning-off.

2,048 questions Binary acceptability Native German ScaLA-de (via EuroEval) ↗
# Model Score Latency Tok in Answer tok Cost Date
1 Claude Fable 5 🔒 reasoning
83.6%
5.20s 1,661.5 28 $44.77 2026-07-06 View result →
2 Claude Opus 4.8
82.6%
2.00s 1,684.8 4 $17.46 2026-06-13 View result →
3 Gemini 3.8 Flash 🔒 reasoning
80.6%
3.45s 777.7 1.5 $4.70 2026-09-04 View result →
4 Gemini 3.1 Pro 🔒 reasoning
80.4%
3.71s 777.9 1.5 $10.37 2026-06-09 View result →
5 GPT-5.5 🔒 reasoning
79.7%
1.52s 913.2 6 $11.05 2026-06-13 View result →
6 Gemini 3.5 Flash
79.4%
0.79s 802.7 1.5 $2.49 2026-06-09 View result →
7 Gemini 3.7 Flash 🔒 reasoning
79.1%
2.18s 777.7 1.5 $1.60 2026-08-14 View result →
8 Gemini 3.6 Flash
77.6%
0.67s 777.7 1.5 $2.41 2026-07-23 View result →
9 GLM-5.3 Flash 🔒 reasoning
77.2%
2.35s 920.6 3.8 2026-08-30 View result →
10 Gemini 2.5 Flash
77.1%
0.40s 802.7 1.5 $0.50 2026-06-09 View result →
11 Gemini 3 Flash Preview
77.0%
0.94s 777.7 1.5 $0.81 2026-06-25 View result →
12 DeepSeek V4 Flash 0731
76.1%
0.76s 949.8 2.5 $0.15 2026-08-11 View result →
13 DeepSeek V4 Pro 0813
75.6%
1.21s 949.8 1.5 $0.85 2026-08-13 View result →
14 GPT-5.6 Terra
75.1%
1.02s 895.5 5.5 $1.91 2026-08-12 View result →
15 Gemini 3.5 Flash-Lite
75.1%
0.50s 777.7 1.5 $0.49 2026-08-30 View result →
16 DeepSeek V4 Flash
74.9%
0.53s 949.8 2.5 $0.27 2026-06-09 View result →
17 MiMo V2.5 Pro
74.8%
0.57s 1,301 2.5 $2.33 2026-06-09 View result →
18 Gemini 3.1 Flash-Lite
74.2%
0.46s 802.7 1.5 $0.83 2026-06-08 View result →
19 DeepSeek V4 Flash 0423
74.0%
0.63s 949.8 2.6 $0.17 2026-08-11 View result →
20 Tencent HY4-Preview
74.0%
1.56s 1,063.5 2.6 2026-08-30 View result →
21 DeepSeek V4 Pro
73.5%
1.41s 949.8 1.6 $3.12 2026-06-09 View result →
22 GPT-5.6 Luna
72.2%
0.85s 895.5 5.6 $0.19 2026-08-13 View result →
23 Qwen3.6 27B (BF16)
71.7%
1.86s 928.9 2.6 2026-08-15 View result →
24 Gemma 4 31B
71.0%
1.28s 910.7 2.5 $0.23 2026-06-11 View result →
25 GLM-5.1
69.9%
1.69s 901.6 2.5 $2.57 2026-06-09 View result →
26 Qwen3.7 Max
69.8%
1.02s 928.9 1.6 $2.39 2026-06-09 View result →
27 Tencent HY3
69.7%
1.42s 1,028.2 2.6 2026-08-30 View result →
28 Qwen3.8 27B (BF16)
66.9%
2.13s 976.9 2.6 2026-08-15 View result →
29 Gemma 4 26B A4B
66.3%
0.46s 910.7 2.6 $0.28 2026-06-08 View result →
30 Qwen3.8 27B (FP8)
66.3%
2.91s 976.9 2.6 2026-08-15 View result →
31 Qwen3.6 35B-A3B
64.2%
0.51s 928.9 2.6 $0.29 2026-06-09 View result →
32 Gemma 4 12B
64.1%
27.85s 910.7 2.6 2026-06-11 View result →
33 Qwen3 14B
64.0%
15.71s 1,060 2.6 2026-06-11 View result →
34 Tencent HY3-Preview
62.5%
2.48s 1,027.6 2.6 $0.13 2026-06-09 View result →
35 Qwen3.5 9B
62.4%
0.82s 928.9 2.5 $0.03 2026-06-10 View result →
36 MiMo V2.5
53.4%
0.74s 1,297 2.7 $0.32 2026-08-11 View result →

TPS is decode speed — output tokens per second after the first token; higher is faster. TTFT is time to first token; lower is snappier. A snapshot, not a constant.

# Model TPS TTFT
1 gpt-oss-120b 🔒 1,641 0.31s
2 Gemini 3.1 Flash-Lite 223 0.61s
3 Gemini 3.5 Flash 181 0.75s
4 Gemini 3.5 Flash-Lite 181 0.71s
5 Gemini 3.6 Flash 166 0.64s
6 Qwen3.6 35B-A3B 159 0.62s
7 Gemini 2.5 Flash 159 0.41s
8 Gemini 3 Flash Preview 147 0.98s
9 Claude Haiku 4.5 131 0.79s
10 DeepSeek V4 Flash 115 0.92s
11 Tencent HY3-Preview 107 2.69s
12 Gemma 4 26B A4B 46 1.16s
13 GLM-5.1 32 0.65s
14 Qwen3 14B 17 0.36s

Quality vs. speed

Average primary score across completed benchmarks against decode speed (output tokens per second). Up and to the right is better — smarter and faster. The gold line is the speed frontier: the best score available at each speed.

55%60%65%70%75%80%85%50100150200300500750100015002000Output speed (tokens / sec)Avg primary scorefaster & smarter — better ↗Gemini 3.5 Flash · 79.1% · 181 tok/sGemini 3.5 FlashGemini 3.1 Flash-Lite · 76.1% · 223 tok/sGemini 3.1 Flash-LiteGemini 3.5 Flash-Lite · 78.4% · 181 tok/sGemini 3.5 Flash-LiteGemini 3.6 Flash · 78.2% · 166 tok/sGemini 3.6 FlashGemini 3 Flash Preview · 77.1% · 147 tok/sGemini 3 Flash PreviewGemini 2.5 Flash · 75.6% · 159 tok/sGemini 2.5 FlashDeepSeek V4 Flash · 74.4% · 115 tok/sDeepSeek V4 FlashClaude Haiku 4.5 · 71.8% · 131 tok/s · partial coverageClaude Haiku 4.5 *Qwen3.6 35B-A3B · 70.2% · 159 tok/s · partial coverageQwen3.6 35B-A3B *gpt-oss-120b · 66.2% · 1641 tok/s · reasoning locked · partial coveragegpt-oss-120b 🔒 *Gemma 4 26B A4B · 65.0% · 46 tok/sGemma 4 26B A4BQwen3 14B · 65.0% · 17 tok/sQwen3 14BTencent HY3-Preview · 64.2% · 107 tok/s · partial coverageTencent HY3-Preview *GLM-5.1 · 61.8% · 32 tok/s · partial coverageGLM-5.1 *

* scored on fewer benchmarks, so its average is not directly comparable to full-coverage models.

🔒 reasoning can't be disabled — its decode speed includes forced reasoning tokens, so it isn't directly comparable to the reasoning-off models.