German Artificial Analytics
a PeerBench project
German-language LLM leaderboard

Which LLM performs best in German?

Independent German-language benchmarks for frontier and fast models — quality, latency, and cost, ranked for real deployment decisions.

Recommended picks

Start here if you need a decision.

Compare models
Best overall

Gemini 3.1 Pro

Highest German aggregate score

75.4 German score
Google 6/7 benchmarks $5.92/1k
Best open weights

DeepSeek V4 Pro

Strongest open-weight model in the ranking

51.9 German score
DeepSeek 7/7 benchmarks $0.451/1k
Fast strong model

Gemini 3.5 Flash

High German score with measured fast throughput

64.3 German score
Google 7/7 benchmarks $3.72/1k 181 tok/s
Best value

Gemini 3.6 Flash

Most German score per listed token price

61.5 German score
Google 7/7 benchmarks $2.81/1k 166 tok/s
Scope
German-language only

No coding, vision, multimodal, or generic English arena score in the public aggregate.

Evidence
31 models · 7 benchmarks

Coverage is shown per model; partial averages are marked instead of hidden.

Benchmarks
Native + translated German

INCLUDE, GermEval, SB10K, ScaLA, MMLU-ProX, MMMLU, and MuSR.

Freshness
Updated 2026-07-23

Published rows are gated before they reach the static site.

Overall German ranking

Official German-language score from German benchmark evidence only. Partial coverage is shown in every row.

Coverage
Weight
Price
Speed
#
Model
Score
Cost
Speed
1
Google · Closed · $5.92/1k · —
75.4
$5.92
/1k questions
tok/s
2
Fable 5 5/7
Anthropic · Closed · $31.45/1k · —
73.2
$31.45
/1k questions
tok/s
3
GPT-5.5 7/7
OpenAI · Closed · $10.28/1k · —
66.6
$10.28
/1k questions
tok/s
4
Opus 4.8 6/7
Anthropic · Closed · $12.16/1k · —
65.8
$12.16
/1k questions
tok/s
5
Google · Closed · $3.72/1k · 181 tok/s
64.3
$3.72
/1k questions
181
tok/s
6
Google · Closed · $2.81/1k · 166 tok/s
61.5
$2.81
/1k questions
166
tok/s
7
Google · Closed · $1.13/1k · 147 tok/s
57.3
$1.13
/1k questions
147
tok/s
8
Google · Closed · $0.582/1k · 223 tok/s
55.6
$0.582
/1k questions
223
tok/s
9
Google · Closed · $0.900/1k · 181 tok/s
54.7
$0.900
/1k questions
181
tok/s
10
Alibaba · Closed · $1.17/1k · —
53.2
$1.17
/1k questions
tok/s
11
DeepSeek · Open weights · $0.451/1k · —
51.9
$0.451
/1k questions
tok/s
12
Google · Open weights · $0.223/1k · —
51.2
$0.223
/1k questions
tok/s
13
Google · Closed · $1.38/1k · 159 tok/s
51.2
$1.38
/1k questions
159
tok/s
14
DeepSeek · Open weights · $0.091/1k · 115 tok/s
50.9
$0.091
/1k questions
115
tok/s
15
Xiaomi · Open weights · $1.00/1k · —
47.7
$1.00
/1k questions
tok/s
16
Alibaba · Open weights · $0.743/1k · 159 tok/s
45.8
$0.743
/1k questions
159
tok/s
17
Google · Open weights · $0.174/1k · 46 tok/s
44.1
$0.174
/1k questions
46
tok/s
18
Google · Open weights · —/1k · —
43.9
/1k questions
tok/s
19
Anthropic · Closed · $3.25/1k · 131 tok/s
43.8
$3.25
/1k questions
131
tok/s
20
Tencent · Open weights · $0.139/1k · 107 tok/s
40.9
$0.139
/1k questions
107
tok/s
21
GLM-5.1 5/7
Z.ai · Open weights · $1.12/1k · 32 tok/s
40.5
$1.12
/1k questions
32
tok/s
22
Alibaba · Open weights · $0.176/1k · —
36.7
$0.176
/1k questions
tok/s
23
Qwen3 14B 7/7
Alibaba · Open weights · —/1k · 17 tok/s
33.9
/1k questions
17
tok/s

Showing 23 models that ran at least 3 of 7 German-language benchmarks (8 excluded for thin coverage). Cost is effective benchmark cost per 1,000 questions; speed is output tokens/sec.

German tradeoffs

Readable picks for value, speed, and German benchmark balance.

Best value

Best price/performance among models within 18 points of the German-score leader.

strong value
1 Gemini 3.6 Flash 7/7 benches
61.5 $2.81/1k 22
2 Gemini 3.5 Flash 7/7 benches
64.3 $3.72/1k 17
3 Gemini 3.1 Pro 6/7 benches
75.4 $5.92/1k 13
4 GPT-5.5 7/7 benches
66.6 $10.28/1k 6
5 Opus 4.8 6/7 benches
65.8 $12.16/1k 5
6 Fable 5 5/7 benches
73.2 $31.45/1k 2

Fast + strong

Fastest measured models that stay within 18 points of the German-score leader.

tok/s rank
1 Gemini 3.5 Flash TTFT 0.75s
64.3 181 tok/s 181
2 Gemini 3.6 Flash TTFT 0.64s
61.5 166 tok/s 166

Native German vs translated German

Dumbbell view: left/right dots show each model's native-German and translated-benchmark averages.

largest gaps
GLM-5.1 -28.2 pp native gap
Native 53.2% Translated 81.4%
Gemma 4 26B A4B -24.9 pp native gap
Native 57.0% Translated 81.9%
Tencent HY3-Preview -16.5 pp native gap
Native 61.9% Translated 78.4%
Qwen3.6 35B-A3B -15.2 pp native gap
Native 67.7% Translated 82.9%
Qwen3.5 9B -13.9 pp native gap
Native 63.8% Translated 77.7%
Gemma 4 31B -13.3 pp native gap
Native 70.8% Translated 84.1%
GPT-5.5 -12.8 pp native gap
Native 75.7% Translated 88.5%
Opus 4.8 -12.3 pp native gap
Native 76.2% Translated 88.5%
Gemini 3.1 Flash-Lite -12.0 pp native gap
Native 72.4% Translated 84.5%
DeepSeek V4 Pro -11.8 pp native gap
Native 71.5% Translated 83.3%

Partial coverage is shown as `n/7`; 🔒 means reasoning is provider-locked. Scores use German-language benchmark evidence only.

GermEval — German NER

Native German named-entity recognition — identify persons, locations, organisations and misc entities in German text, emitted as JSON. Scored with seqeval micro-F1 excluding the noisy MISC class. Run reasoning-off.

1,024 questions Named-entity recognition Native German GermEval (via EuroEval) ↗
# Model Score Latency Tok in Answer tok Cost Date
1 Fable 5 🔒 reasoning
87.7%
5.52s 2,083.6 76.1 $30.07 2026-07-06
2 Gemini 3.1 Pro 🔒 reasoning
87.3%
4.86s 1,059.7 29.5 $10.04 2026-06-09
3 Gemini 3.6 Flash
86.2%
1.04s 1,060 38.2 $1.92 2026-07-23
4 Gemini 3.5 Flash
86.0%
0.81s 1,077 24.6 $1.88 2026-06-09
5 GPT-5.5 🔒 reasoning
85.9%
1.86s 1,140.7 29.1 $7.08 2026-06-13
6 Opus 4.8
85.8%
2.12s 2,110.5 42 $11.88 2026-06-13
7 Gemini 3.5 Flash-Lite
85.4%
0.70s 1,060 46.3 $0.44 2026-07-23
8 Gemini 3 Flash Preview
84.7%
1.20s 1,060 24.9 $0.62 2026-06-25
9 Gemini 3.1 Flash-Lite
83.2%
0.52s 1,077 25.7 $0.63 2026-06-08
10 Gemma 4 31B
82.6%
11.92s 1,153 27.8 $0.15 2026-06-11
11 DeepSeek V4 Pro
82.2%
1.74s 1,181.2 26.3 $2.02 2026-06-09
12 Qwen3.7 Max
82.1%
1.64s 1,154.3 25.2 $1.57 2026-06-09
13 MiMo V2.5 Pro
81.9%
1.09s 1,533.5 26.4 $1.49 2026-06-09
14 Gemma 4 26B A4B
81.8%
1.14s 1,153 30.3 $0.19 2026-06-08
15 Gemini 2.5 Flash
81.8%
0.52s 1,077 24.5 $0.39 2026-06-09
16 DeepSeek V4 Flash
80.2%
0.77s 1,181.2 26.8 $0.18 2026-06-09
17 Qwen3.6 35B-A3B
79.9%
0.67s 1,154.3 27.3 $0.21 2026-06-09
18 Gemma 4 12B
79.6%
43.80s 1,153 28.1 2026-06-11
19 Tencent HY3-Preview
77.3%
2.50s 1,313.7 35.4 $0.09 2026-06-09
20 Qwen3 14B
73.0%
21.12s 1,292.5 28.5 2026-06-11
21 Qwen3.5 9B
72.6%
11.77s 1,154.3 26.9 $0.12 2026-06-11
22 GLM-5.1
52.2%
2.20s 1,127.6 26.5 $1.71 2026-06-09

INCLUDE — German

Native German exam and licensing questions covering region-specific knowledge — history, law, civics and culture. Written by humans in German, not translated.

139 questions 4-option multiple choice Native German CohereLabs/include-base-44 ↗
# Model Score Latency Tok in Answer tok Cost Date
1 Gemini 3.1 Pro 🔒 reasoning
77.7%
4.24s 107.1 128.9 $0.72 2026-06-10
2 GPT-5.5 🔒 reasoning
74.8%
3.99s 128.2 50.2 $0.70 2026-06-12
3 Gemini 3.6 Flash
74.8%
1.37s 105.8 168.7 $0.19 2026-07-23
4 Gemini 3.5 Flash
74.1%
1.72s 107.1 205.4 $0.28 2026-06-10
5 MiniMax M2.7 🔒 reasoning
72.7%
12.26s 151.1 -23.7 $0.29
6 Fable 5 🔒 reasoning
72.7%
7.72s 228.2 437 $3.33 2026-07-05
7 Sonnet 5
72.7%
4.92s 229.9 263.5 $0.64 2026-07-21
8 Gemini 3.1 Flash-Lite
71.9%
0.89s 107.1 151 $0.04 2026-06-10
9 Opus 4.8
71.9%
5.77s 256.2 343.7 $1.32 2026-06-12
10 Gemini 3 Flash Preview
71.9%
2.37s 106.4 170.5 $0.08 2026-06-25
11 Qwen3.7 Max
71.2%
2026-06-08
12 Gemini 2.5 Flash
70.5%
1.61s 105.8 193.6 $0.07 2026-06-02
13 DeepSeek V4 Pro
70.5%
3.44s 118.7 121.3 $0.08 2026-06-08
14 DeepSeek V4 Flash
70.5%
2.53s 118.626 111.288 $0.0066 2026-06-08
15 Gemini 3.5 Flash-Lite
70.5%
1.44s 106.2 235.4 $0.09 2026-07-23
16 Kimi K2.6
69.8%
13.65s 138.7 527.1 $0.58 2026-06-08
17 Tencent HY3-Preview
69.1%
5.48s 143.9 222.4 $0.0086 2026-06-02
18 Gemma 4 12B
69.1%
27.64s 123.1 221.9 2026-06-11
19 Claude Haiku 4.5
68.3%
3.45s 154.6 347.4 $0.26 2026-06-02
20 Qwen3.6 35B-A3B
68.3%
2.72s 118.4 545.1 $0.08 2026-06-02
21 Gemma 4 31B
68.3%
4.56s 119.1 205.4 $0.01 2026-06-11
22 MiMo V2.5 Pro
67.6%
2026-06-08
23 GLM-5.1
67.6%
2026-06-08
24 gpt-oss-120b 🔒 reasoning
66.2%
1.71s 154.1 21.4 $0.0078 2026-05-29
25 Gemma 4 26B A4B
64.7%
3.71s 118.7 275.8 $0.03 2026-06-03
26 Qwen3.5 9B
64.7%
10.95s 122.4 378 $0.0096 2026-06-11
27 grok-4.3
63.3%
0.60s 233.8 20.5 $0.04 2026-06-03
28 Qwen3 14B
63.3%
3.32s 135.4 46 2026-06-11
29 Ministral 14B
58.3%
0.41s 109.1 79.8 $0.01 2026-06-03
30 gemma-3-12b-it
53.2%
5.00s 114.8 160.5 $0.0073 2026-06-03

MMLU-ProX — German

Hard academic questions across 14 subjects — STEM, law, health, economics, philosophy and more. Professionally translated to German, with up to ten answer options per question.

11,759 questions 10-option multiple choice Professional translation li-lab/MMLU-ProX ↗
# Model Score Latency Tok in Answer tok Cost Date
1 GPT-5.5 🔒 reasoning
87.2%
29.48s 1,594.7 145.5 $193.51 2026-07-01
2 Gemini 3.5 Flash
86.5%
2.23s 1,664.9 353.3 $66.76 2026-05-26
3 Gemini 3 Flash Preview
86.3%
3.26s 1,664.9 311.5 $20.78 2026-06-25
4 Gemini 3.6 Flash
86.0%
1.84s 1,664.9 259.2 $52.23 2026-07-23
5 Gemini 3.1 Flash-Lite
82.2%
1.24s 1,666.1 304.5 $10.21 2026-06-03
6 Gemma 4 31B
82.1%
64.35s 1,702.9 502.5 $4.53 2026-06-11
7 DeepSeek V4 Pro
80.8%
3.00s 288.1 218.9 $3.71 2026-06-16
8 Gemini 3.5 Flash-Lite
80.2%
1.57s 1,664.9 307.2 $14.90 2026-07-23
9 Qwen3.6 35B-A3B
80.0%
4.67s 1,689.6 984 $13.36 2026-05-31
10 Gemini 2.5 Flash
79.6%
2.69s 1,664.9 786 $28.89 2026-05-26
11 Gemma 4 26B A4B
78.2%
7.77s 1,700.6 662.1 $4.64 2026-06-05
12 Claude Haiku 4.5
75.3%
3.76s 2,262 433.7 $52.10 2026-06-03
13 DeepSeek V4 Flash
74.9%
1.98s 1,718.7 162.2 $1.04 2026-06-16
14 Qwen3.5 9B
73.4%
40.41s 1,693.6 931.1 $3.63 2026-06-11
15 Tencent HY3-Preview
73.2%
6.14s 1,910.2 586.3 $2.76 2026-05-27
16 Gemma 4 12B
73.1%
92.81s 1,705.9 493.3 2026-06-11
17 Gemini 2.5 Flash-Lite
71.2%
2.48s 1,665.1 1,493.5 $8.95 2026-06-03
18 Qwen3 14B
64.1%
29.26s 1,887.3 413.7 2026-06-11
Show per-subject breakdown (252)
Subject Model Score
biology Gemini 3.5 Flash 92.9%
biology Gemini 3 Flash Preview 92.6%
biology Gemini 3.6 Flash 92.6%
biology GPT-5.5 92.3%
biology Gemma 4 31B 91.2%
biology DeepSeek V4 Pro 90.2%
biology Qwen3.6 35B-A3B 89.7%
biology Gemini 3.1 Flash-Lite 89.4%
biology Gemini 2.5 Flash 88.8%
biology DeepSeek V4 Flash 88.4%
biology Gemini 3.5 Flash-Lite 88.1%
biology Gemma 4 26B A4B 87.3%
biology Gemini 2.5 Flash-Lite 86.8%
biology Claude Haiku 4.5 85.9%
biology Tencent HY3-Preview 85.8%
business GPT-5.5 91.1%
business Gemini 3.5 Flash 89.6%
business Gemini 3 Flash Preview 89.5%
business Gemini 3.6 Flash 88.2%
business Gemma 4 31B 87.5%
business Gemini 3.1 Flash-Lite 85.3%
business DeepSeek V4 Pro 85.3%
business Gemini 3.5 Flash-Lite 84.4%
business Gemma 4 26B A4B 83.9%
business Gemini 2.5 Flash 83.5%
business Qwen3.6 35B-A3B 83.5%
business Claude Haiku 4.5 79.0%
business Tencent HY3-Preview 78.3%
business Gemini 2.5 Flash-Lite 77.7%
business DeepSeek V4 Flash 76.2%
chemistry GPT-5.5 89.8%
chemistry Gemini 3 Flash Preview 89.2%
chemistry Gemini 3.5 Flash 89.1%
chemistry Gemini 3.6 Flash 88.3%
chemistry Qwen3.6 35B-A3B 88.2%
chemistry Gemini 3.1 Flash-Lite 87.1%
chemistry Gemini 2.5 Flash 86.9%
chemistry Gemma 4 31B 86.7%
chemistry DeepSeek V4 Pro 86.7%
chemistry Gemma 4 26B A4B 85.4%
chemistry Gemini 3.5 Flash-Lite 83.8%
chemistry Claude Haiku 4.5 80.4%
chemistry Gemini 2.5 Flash-Lite 78.4%
chemistry Tencent HY3-Preview 75.6%
chemistry DeepSeek V4 Flash 73.3%
computer science GPT-5.5 90.0%
computer science Gemini 3.1 Flash-Lite 88.8%
computer science Gemini 3.6 Flash 88.5%
computer science Gemini 3.5 Flash 87.6%
computer science Gemini 3 Flash Preview 86.6%
computer science DeepSeek V4 Flash 85.9%
computer science Gemma 4 31B 85.9%
computer science DeepSeek V4 Pro 85.4%
computer science Gemini 2.5 Flash 85.1%
computer science Qwen3.6 35B-A3B 85.1%
computer science Gemma 4 26B A4B 83.9%
computer science Claude Haiku 4.5 83.4%
computer science Gemini 3.5 Flash-Lite 83.2%
computer science Gemini 2.5 Flash-Lite 77.8%
computer science Tencent HY3-Preview 76.8%
economics GPT-5.5 89.2%
economics Gemini 3.5 Flash 89.1%
economics Gemini 3 Flash Preview 88.7%
economics Gemini 3.6 Flash 88.7%
economics Gemini 3.1 Flash-Lite 87.3%
economics Gemma 4 31B 86.8%
economics Gemini 2.5 Flash 86.4%
economics Gemini 3.5 Flash-Lite 85.9%
economics Qwen3.6 35B-A3B 85.3%
economics DeepSeek V4 Pro 84.1%
economics Gemma 4 26B A4B 82.6%
economics Tencent HY3-Preview 82.3%
economics Claude Haiku 4.5 81.5%
economics Gemini 2.5 Flash-Lite 79.1%
economics DeepSeek V4 Flash 75.0%
engineering GPT-5.5 83.3%
engineering Gemini 3.5 Flash 82.5%
engineering Gemini 3 Flash Preview 81.2%
engineering Gemini 3.6 Flash 80.5%
engineering Gemini 3.1 Flash-Lite 77.8%
engineering Qwen3.6 35B-A3B 77.5%
engineering Gemma 4 31B 77.1%
engineering DeepSeek V4 Pro 74.2%
engineering Gemma 4 26B A4B 73.5%
engineering Gemini 3.5 Flash-Lite 72.8%
engineering Gemini 2.5 Flash 71.6%
engineering Claude Haiku 4.5 64.7%
engineering Tencent HY3-Preview 64.4%
engineering DeepSeek V4 Flash 64.1%
engineering Gemini 2.5 Flash-Lite 58.4%
health GPT-5.5 81.8%
health Gemini 3 Flash Preview 80.1%
health Gemini 3.6 Flash 79.3%
health Gemini 3.5 Flash 78.6%
health Gemini 3.1 Flash-Lite 74.8%
health DeepSeek V4 Pro 74.2%
health Gemma 4 31B 74.1%
health Gemini 3.5 Flash-Lite 73.9%
health Tencent HY3-Preview 73.7%
health Claude Haiku 4.5 72.9%
health Qwen3.6 35B-A3B 72.8%
health Gemini 2.5 Flash 72.5%
health DeepSeek V4 Flash 71.3%
health Gemma 4 26B A4B 71.2%
health Gemini 2.5 Flash-Lite 66.5%
history Gemini 3 Flash Preview 80.8%
history Gemini 3.5 Flash 80.6%
history Gemini 3.6 Flash 80.1%
history GPT-5.5 77.7%
history Gemini 3.5 Flash-Lite 75.1%
history Gemini 3.1 Flash-Lite 74.5%
history Gemma 4 31B 74.5%
history DeepSeek V4 Pro 71.9%
history Tencent HY3-Preview 71.4%
history Gemini 2.5 Flash 70.9%
history Qwen3.6 35B-A3B 69.0%
history DeepSeek V4 Flash 68.2%
history Gemma 4 26B A4B 66.7%
history Claude Haiku 4.5 65.4%
history Gemini 2.5 Flash-Lite 59.3%
law GPT-5.5 74.8%
law Gemini 3 Flash Preview 73.2%
law Gemini 3.6 Flash 73.2%
law Gemini 3.5 Flash 72.8%
law Gemini 3.1 Flash-Lite 62.8%
law Gemma 4 31B 59.0%
law Gemini 3.5 Flash-Lite 58.2%
law DeepSeek V4 Pro 55.3%
law Gemini 2.5 Flash 54.0%
law Qwen3.6 35B-A3B 52.8%
law Gemma 4 26B A4B 50.7%
law Claude Haiku 4.5 47.4%
law DeepSeek V4 Flash 46.1%
law Tencent HY3-Preview 45.8%
law Gemini 2.5 Flash-Lite 41.6%
math Gemini 3.5 Flash 94.6%
math GPT-5.5 94.6%
math Gemini 3 Flash Preview 94.1%
math Gemini 3.6 Flash 93.9%
math Gemma 4 31B 93.5%
math Qwen3.6 35B-A3B 92.7%
math Gemma 4 26B A4B 92.0%
math DeepSeek V4 Pro 90.9%
math Gemini 2.5 Flash 90.5%
math Gemini 3.1 Flash-Lite 90.2%
math Gemini 3.5 Flash-Lite 89.9%
math DeepSeek V4 Flash 88.7%
math Claude Haiku 4.5 86.8%
math Tencent HY3-Preview 86.5%
math Gemini 2.5 Flash-Lite 82.8%
nothink biology Gemma 4 12B 87.0%
nothink biology Qwen3.5 9B 85.6%
nothink biology Qwen3 14B 81.3%
nothink business Gemma 4 12B 79.8%
nothink business Qwen3.5 9B 78.1%
nothink business Qwen3 14B 70.3%
nothink chemistry Qwen3.5 9B 85.0%
nothink chemistry Gemma 4 12B 79.9%
nothink chemistry Qwen3 14B 72.2%
nothink computer science Gemma 4 12B 79.3%
nothink computer science Qwen3.5 9B 78.0%
nothink computer science Qwen3 14B 71.5%
nothink economics Qwen3.5 9B 78.9%
nothink economics Gemma 4 12B 78.7%
nothink economics Qwen3 14B 71.1%
nothink engineering Qwen3.5 9B 66.6%
nothink engineering Gemma 4 12B 65.1%
nothink engineering Qwen3 14B 55.9%
nothink health Qwen3.5 9B 65.9%
nothink health Gemma 4 12B 64.5%
nothink health Qwen3 14B 56.8%
nothink history Qwen3.5 9B 60.4%
nothink history Gemma 4 12B 57.2%
nothink history Qwen3 14B 48.8%
nothink law Gemma 4 12B 42.8%
nothink law Qwen3.5 9B 36.7%
nothink law Qwen3 14B 27.4%
nothink math Qwen3.5 9B 90.2%
nothink math Gemma 4 12B 90.1%
nothink math Qwen3 14B 82.2%
nothink other Qwen3.5 9B 61.6%
nothink other Gemma 4 12B 60.7%
nothink other Qwen3 14B 52.1%
nothink philosophy Gemma 4 12B 60.3%
nothink philosophy Qwen3.5 9B 58.1%
nothink philosophy Qwen3 14B 49.3%
nothink physics Qwen3.5 9B 84.7%
nothink physics Gemma 4 12B 81.9%
nothink physics Qwen3 14B 73.1%
nothink psychology Gemma 4 12B 76.3%
nothink psychology Qwen3.5 9B 74.2%
nothink psychology Qwen3 14B 66.0%
other GPT-5.5 84.3%
other Gemini 3 Flash Preview 82.6%
other Gemini 3.5 Flash 82.4%
other Gemini 3.6 Flash 81.5%
other DeepSeek V4 Pro 77.3%
other Gemini 3.1 Flash-Lite 76.7%
other Gemini 3.5 Flash-Lite 75.3%
other Gemma 4 31B 73.7%
other Gemini 2.5 Flash 73.5%
other Tencent HY3-Preview 72.1%
other Qwen3.6 35B-A3B 70.8%
other Claude Haiku 4.5 69.5%
other DeepSeek V4 Flash 69.4%
other Gemma 4 26B A4B 67.5%
other Gemini 2.5 Flash-Lite 64.9%
philosophy GPT-5.5 85.2%
philosophy Gemini 3.5 Flash 83.2%
philosophy Gemini 3 Flash Preview 82.4%
philosophy Gemini 3.6 Flash 82.4%
philosophy Gemini 3.1 Flash-Lite 75.2%
philosophy Gemma 4 31B 74.3%
philosophy DeepSeek V4 Pro 73.7%
philosophy Gemini 3.5 Flash-Lite 72.9%
philosophy DeepSeek V4 Flash 71.9%
philosophy Gemini 2.5 Flash 70.9%
philosophy Qwen3.6 35B-A3B 69.5%
philosophy Gemma 4 26B A4B 69.1%
philosophy Tencent HY3-Preview 69.1%
philosophy Claude Haiku 4.5 67.9%
philosophy Gemini 2.5 Flash-Lite 60.9%
physics GPT-5.5 90.8%
physics Gemini 3.5 Flash 90.4%
physics Gemini 3.6 Flash 90.4%
physics Gemini 3 Flash Preview 90.3%
physics Gemma 4 31B 89.0%
physics Qwen3.6 35B-A3B 88.0%
physics Gemini 3.1 Flash-Lite 87.6%
physics DeepSeek V4 Pro 86.8%
physics Gemma 4 26B A4B 86.1%
physics Gemini 3.5 Flash-Lite 85.9%
physics Gemini 2.5 Flash 85.5%
physics DeepSeek V4 Flash 84.7%
physics Claude Haiku 4.5 81.3%
physics Gemini 2.5 Flash-Lite 76.6%
physics Tencent HY3-Preview 64.7%
psychology Gemini 3.5 Flash 87.8%
psychology Gemini 3 Flash Preview 87.8%
psychology Gemini 3.6 Flash 87.3%
psychology GPT-5.5 86.3%
psychology Gemini 3.5 Flash-Lite 84.0%
psychology Gemma 4 31B 83.6%
psychology Gemini 3.1 Flash-Lite 83.3%
psychology DeepSeek V4 Pro 83.3%
psychology Gemini 2.5 Flash 82.6%
psychology DeepSeek V4 Flash 80.8%
psychology Tencent HY3-Preview 80.5%
psychology Claude Haiku 4.5 78.9%
psychology Qwen3.6 35B-A3B 78.7%
psychology Gemma 4 26B A4B 78.6%
psychology Gemini 2.5 Flash-Lite 75.1%

MMMLU — German

OpenAI's multilingual MMLU, German split — general knowledge spanning STEM, the humanities, social sciences and other domains. Professionally translated to German.

14,042 questions 4-option multiple choice Professional translation openai/MMMLU ↗
# Model Score Latency Tok in Answer tok Cost Date
1 Gemini 3.1 Pro 🔒 reasoning
92.7%
4.85s 172.5 158.5 $76.03 2026-07-13
2 GPT-5.5 🔒 reasoning
92.1%
4.53s 198.9 74.9 $83.60 2026-06-16
3 Opus 4.8
90.7%
7.20s 383.2 404.2 $167.48 2026-06-25
4 Gemini 3 Flash Preview
90.1%
2.80s 172 220.6 $10.55 2026-06-25
5 Gemini 3.6 Flash
89.5%
1.82s 172.5 198 $24.30 2026-07-23
6 Gemini 3.5 Flash
89.3%
2.05s 172 255.8 $35.95 2026-05-26
7 Gemini 3.1 Flash-Lite
86.8%
1.06s 173 210.7 $5.04 2026-06-03
8 Gemma 4 31B
86.6%
12.07s 185 251.5 $1.58 2026-06-11
9 Gemini 3.5 Flash-Lite
86.0%
1.51s 172.5 273.6 $10.25 2026-07-23
10 Qwen3.6 35B-A3B
85.8%
4.39s 182.9 561.4 $8.27 2026-06-01
11 DeepSeek V4 Flash
85.2%
2.82s 190.8 139.1 $0.91 2026-06-23
12 Gemini 2.5 Flash
84.7%
1.48s 172 275.5 $10.40 2026-05-26
13 DeepSeek V4 Pro
84.0%
3.51s 190.8 139.3 $2.83 2026-06-22
14 Gemma 4 26B A4B
83.7%
7.77s 112.3 221.7 $0.05 2026-06-05
15 Tencent HY3-Preview
83.7%
5.55s 223.4 280.9 $1.14 2026-05-28
16 Claude Haiku 4.5
83.1%
4.20s 277.7 397.6 $31.90 2026-06-02
17 Gemini 2.5 Flash-Lite
79.6%
1.72s 173 506 $3.06 2026-06-03
18 Gemma 4 12B
79.0%
47.38s 189 261.2 2026-06-11
19 Qwen3.5 9B
78.8%
13.81s 186.9 408.3 $1.12 2026-06-11
20 Qwen3 14B
73.4%
6.37s 212 87.9 2026-06-11

MuSR — German

Multi-step soft reasoning over long narrative contexts — murder mysteries, object placement and team allocation. Requires chaining clues across several paragraphs to reach the correct answer. Translated to German from the original English MuSR benchmark.

564 questions 2–5 option multiple choice Professional translation zayne-sprague/MuSR ↗
# Model Score Latency Tok in Answer tok Cost Date
1 Fable 5 🔒 reasoning
89.5%
20.28s 3,008.8 1,338.2 $54.71 2026-07-05
2 Gemini 3.1 Pro 🔒 reasoning
88.1%
12.14s 1,464.9 510.8 $10.49 2026-06-09
3 GPT-5.5 🔒 reasoning
86.3%
13.58s 1,449.5 414.8 $14.53 2026-06-13
4 Opus 4.8
86.3%
13.87s 3,039.8 1,075.6 $23.74 2026-06-13
5 DeepSeek V4 Pro
85.3%
12.60s 1,657.8 719.7 $0.74 2026-06-23
6 Gemini 3.1 Flash-Lite
84.4%
2.95s 1,467.3 568.5 $0.73 2026-06-09
7 Gemini 3.5 Flash
84.2%
5.79s 1,464.9 850.2 $5.55 2026-06-09
8 Gemma 4 26B A4B
83.7%
2026-06-10
9 DeepSeek V4 Flash
83.5%
10.18s 1,659.8 798.9 $0.26 2026-06-10
10 Gemma 4 31B
83.5%
23.30s 1,477.9 650.1 $0.23 2026-06-11
11 Gemini 2.5 Flash
83.3%
6.28s 1,464.9 1,077.4 $1.77 2026-06-09
12 Gemini 3.5 Flash-Lite
81.7%
2.86s 1,464.9 649.3 $1.16 2026-07-23
13 Gemma 4 12B
81.6%
195.50s 1,477.9 707.4 2026-06-11
14 GLM-5.1
81.4%
2026-06-10
15 MiMo V2.5 Pro
80.9%
2026-06-10
16 Qwen3.5 9B
80.9%
61.14s 1,469.2 1,629.7 $0.22 2026-06-11
17 Gemini 3.6 Flash
80.3%
4.01s 1,464.9 663.7 $4.05 2026-07-23
18 Gemini 3 Flash Preview
79.3%
5.41s 1,464.9 622.9 $1.47 2026-06-25
19 Qwen3 14B
69.7%
263.43s 1,728.9 1,270.4 2026-06-11
Show per-subject breakdown (57)
Subject Model Score
murder mystery Fable 5 92.8%
murder mystery GPT-5.5 90.0%
murder mystery Gemini 3.1 Pro 89.2%
murder mystery Gemini 2.5 Flash 87.6%
murder mystery Opus 4.8 87.6%
murder mystery Gemma 4 31B 86.4%
murder mystery DeepSeek V4 Pro 86.4%
murder mystery DeepSeek V4 Flash 85.6%
murder mystery Gemma 4 12B 85.6%
murder mystery Gemma 4 26B A4B 85.2%
murder mystery Gemini 3.5 Flash-Lite 85.2%
murder mystery Gemini 3.1 Flash-Lite 84.8%
murder mystery Gemini 3.5 Flash 84.0%
murder mystery GLM-5.1 80.0%
murder mystery Gemini 3.6 Flash 79.6%
murder mystery MiMo V2.5 Pro 78.8%
murder mystery Gemini 3 Flash Preview 78.8%
murder mystery Qwen3.5 9B 77.6%
murder mystery Qwen3 14B 76.8%
object placements Gemini 3.5 Flash 81.3%
object placements GLM-5.1 81.3%
object placements Fable 5 81.3%
object placements Gemini 3.1 Pro 79.7%
object placements Opus 4.8 79.7%
object placements Gemma 4 12B 79.7%
object placements MiMo V2.5 Pro 76.6%
object placements Gemma 4 31B 76.6%
object placements Gemini 3 Flash Preview 76.6%
object placements Gemini 3.1 Flash-Lite 75.0%
object placements Gemini 2.5 Flash 75.0%
object placements DeepSeek V4 Flash 75.0%
object placements Gemini 3.5 Flash-Lite 75.0%
object placements Gemma 4 26B A4B 73.4%
object placements Qwen3.5 9B 73.4%
object placements Gemini 3.6 Flash 73.4%
object placements Qwen3 14B 71.9%
object placements DeepSeek V4 Pro 70.3%
object placements GPT-5.5 68.8%
team allocation Gemini 3.1 Pro 89.2%
team allocation Fable 5 88.4%
team allocation DeepSeek V4 Pro 88.0%
team allocation GPT-5.5 87.2%
team allocation Opus 4.8 86.8%
team allocation Gemini 3.1 Flash-Lite 86.4%
team allocation Qwen3.5 9B 86.0%
team allocation Gemini 3.5 Flash 85.2%
team allocation DeepSeek V4 Flash 84.8%
team allocation Gemma 4 26B A4B 84.8%
team allocation MiMo V2.5 Pro 84.0%
team allocation GLM-5.1 82.8%
team allocation Gemini 3.6 Flash 82.8%
team allocation Gemma 4 31B 82.4%
team allocation Gemini 2.5 Flash 81.2%
team allocation Gemini 3 Flash Preview 80.4%
team allocation Gemini 3.5 Flash-Lite 80.0%
team allocation Gemma 4 12B 78.0%
team allocation Qwen3 14B 62.0%

SB10K — German sentiment

Native German social-media sentiment classification — positive, neutral or negative. Human-annotated German text, not translated. Run reasoning-off; headline score is MCC on the predicted label.

1,024 questions 3-class sentiment Native German SB10K (via EuroEval) ↗
# Model Score Latency Tok in Answer tok Cost Date
1 Gemini 3.1 Pro 🔒 reasoning
70.1%
3.30s 625.7 1.7 $3.81 2026-06-09
2 Fable 5 🔒 reasoning
67.3%
4.60s 1,352.6 19.3 $18.05 2026-07-06
3 Opus 4.8
64.5%
1.88s 1,372.7 5 $7.16 2026-06-13
4 Gemini 3.6 Flash
64.5%
0.93s 625.7 1.7 $0.97 2026-07-23
5 Gemini 3.5 Flash
64.3%
0.67s 650.6 1.7 $1.02 2026-06-09
6 Gemini 3 Flash Preview
63.6%
1.04s 625.7 1.7 $0.33 2026-06-25
7 Qwen3.7 Max
63.2%
0.99s 771.9 1.7 $0.99 2026-06-09
8 Gemini 3.5 Flash-Lite
62.5%
0.59s 625.7 1.7 $0.20 2026-07-23
9 GPT-5.5 🔒 reasoning
62.5%
1.36s 756.9 5.9 $4.12 2026-06-13
10 DeepSeek V4 Flash
62.4%
0.54s 778.6 2.7 $0.11 2026-06-09
11 Gemma 4 12B
62.2%
21.45s 758.7 2.7 2026-06-11
12 MiMo V2.5 Pro
61.7%
0.45s 1,140.3 2.8 $0.99 2026-06-09
13 Gemini 2.5 Flash
61.6%
0.40s 650.7 1.8 $0.20 2026-06-09
14 Gemma 4 31B
61.0%
5.46s 758.7 2.8 $0.09 2026-06-11
15 Gemini 3.1 Flash-Lite
60.4%
0.46s 650.7 1.8 $0.34 2026-06-08
16 DeepSeek V4 Pro
60.0%
1.42s 778.6 2.1 $1.28 2026-06-09
17 Qwen3.6 35B-A3B
58.4%
0.45s 771.9 2.7 $0.12 2026-06-09
18 Qwen3 14B
56.1%
13.26s 899.3 2.7 2026-06-11
19 Qwen3.5 9B
55.7%
6.96s 771.9 2.6 $0.08 2026-06-11
20 Tencent HY3-Preview
38.9%
2.35s 886.7 2.7 $0.06 2026-06-09
21 GLM-5.1
23.1%
1.88s 765.6 2.6 $1.09 2026-06-09
22 Gemma 4 26B A4B
15.2%
0.70s 758.8 2.8 $0.12 2026-06-08

ScaLA — German acceptability

Native German linguistic acceptability — does the sentence read as grammatical German (ja / nein)? Built from clean vs. minimally-corrupted German sentences. Run reasoning-off.

2,048 questions Binary acceptability Native German ScaLA-de (via EuroEval) ↗
# Model Score Latency Tok in Answer tok Cost Date
1 Fable 5 🔒 reasoning
83.6%
5.20s 1,661.5 28 $44.77 2026-07-06
2 Opus 4.8
82.6%
2.00s 1,684.8 4 $17.46 2026-06-13
3 Gemini 3.1 Pro 🔒 reasoning
80.4%
3.71s 777.9 1.5 $10.37 2026-06-09
4 GPT-5.5 🔒 reasoning
79.7%
1.52s 913.2 6 $11.05 2026-06-13
5 Gemini 3.5 Flash
79.4%
0.79s 802.7 1.5 $2.49 2026-06-09
6 Gemini 3.6 Flash
77.6%
0.67s 777.7 1.5 $2.41 2026-07-23
7 Gemini 2.5 Flash
77.1%
0.40s 802.7 1.5 $0.50 2026-06-09
8 Gemini 3 Flash Preview
77.0%
0.94s 777.7 1.5 $0.81 2026-06-25
9 Gemini 3.5 Flash-Lite
75.1%
0.56s 777.7 1.5 $0.49 2026-07-23
10 DeepSeek V4 Flash
74.9%
0.53s 949.8 2.5 $0.27 2026-06-09
11 MiMo V2.5 Pro
74.8%
0.57s 1,301 2.5 $2.33 2026-06-09
12 Gemini 3.1 Flash-Lite
74.2%
0.46s 802.7 1.5 $0.83 2026-06-08
13 DeepSeek V4 Pro
73.5%
1.41s 949.8 1.6 $3.12 2026-06-09
14 Gemma 4 31B
71.0%
1.28s 910.7 2.5 $0.23 2026-06-11
15 GLM-5.1
69.9%
1.69s 901.6 2.5 $2.57 2026-06-09
16 Qwen3.7 Max
69.8%
1.02s 928.9 1.6 $2.39 2026-06-09
17 Gemma 4 26B A4B
66.3%
0.46s 910.7 2.6 $0.28 2026-06-08
18 Qwen3.6 35B-A3B
64.2%
0.51s 928.9 2.6 $0.29 2026-06-09
19 Gemma 4 12B
64.1%
27.85s 910.7 2.6 2026-06-11
20 Qwen3 14B
64.0%
15.71s 1,060 2.6 2026-06-11
21 Tencent HY3-Preview
62.5%
2.48s 1,027.6 2.6 $0.13 2026-06-09
22 Qwen3.5 9B
62.1%
8.40s 928.9 2.5 $0.19 2026-06-11

TPS is decode speed — output tokens per second after the first token; higher is faster. TTFT is time to first token; lower is snappier. A snapshot, not a constant.

# Model TPS TTFT
1 gpt-oss-120b 🔒 1,641 0.31s
2 Gemini 3.1 Flash-Lite 223 0.61s
3 Gemini 3.5 Flash 181 0.75s
4 Gemini 3.5 Flash-Lite 181 0.71s
5 Gemini 3.6 Flash 166 0.64s
6 Qwen3.6 35B-A3B 159 0.62s
7 Gemini 2.5 Flash 159 0.41s
8 Gemini 3 Flash Preview 147 0.98s
9 Claude Haiku 4.5 131 0.79s
10 DeepSeek V4 Flash 115 0.92s
11 Tencent HY3-Preview 107 2.69s
12 Gemma 4 26B A4B 46 1.16s
13 GLM-5.1 32 0.65s
14 Qwen3 14B 17 0.36s

Quality vs. speed

Average primary score across completed benchmarks against decode speed (output tokens per second). Up and to the right is better — smarter and faster. The gold line is the speed frontier: the best score available at each speed.

55%60%65%70%75%80%85%50100150200300500750100015002000Output speed (tokens / sec)Avg primary scorefaster & smarter — better ↗Gemini 3.5 Flash · 80.6% · 181 tok/sGemini 3.5 FlashGemini 3.1 Flash-Lite · 77.6% · 223 tok/sGemini 3.1 Flash-LiteGemini 3.6 Flash · 79.8% · 166 tok/sGemini 3.6 FlashGemini 3 Flash Preview · 79.0% · 147 tok/sGemini 3 Flash PreviewGemini 3.5 Flash-Lite · 77.4% · 181 tok/sGemini 3.5 Flash-LiteGemini 2.5 Flash · 76.9% · 159 tok/sGemini 2.5 FlashDeepSeek V4 Flash · 75.9% · 115 tok/sDeepSeek V4 FlashClaude Haiku 4.5 · 75.6% · 131 tok/s · partial coverageClaude Haiku 4.5 *Qwen3.6 35B-A3B · 72.8% · 159 tok/s · partial coverageQwen3.6 35B-A3B *gpt-oss-120b · 66.2% · 1641 tok/s · reasoning locked · partial coveragegpt-oss-120b 🔒 *Gemma 4 26B A4B · 67.7% · 46 tok/sGemma 4 26B A4BTencent HY3-Preview · 67.4% · 107 tok/s · partial coverageTencent HY3-Preview *Qwen3 14B · 66.2% · 17 tok/sQwen3 14BGLM-5.1 · 58.8% · 32 tok/s · partial coverageGLM-5.1 *

* scored on fewer benchmarks, so its average is not directly comparable to full-coverage models.

🔒 reasoning can't be disabled — its decode speed includes forced reasoning tokens, so it isn't directly comparable to the reasoning-off models.