German LLM Benchmark · Model Profile

Gemma 4 26B A4B

Google provider-internal run 2026-06-10

67 .7%

avg. German score

#23 of 30 models

−4.5pp below avg.

Benchmark breakdown

GermEval

81.8%

Native German named-entity recognition — identify persons, locations, organisations and misc entities in German text, emitted as JSON. Scored with seqeval micro-F1 excluding the noisy MISC class. Run reasoning-off.

Named-entity recognition native · Native German

via GermEval (via EuroEval) ↗

INCLUDE

64.7%

Native German exam and licensing questions covering region-specific knowledge — history, law, civics and culture. Written by humans in German, not translated.

4-option multiple choice native · Native German

via CohereLabs/include-base-44 ↗

MMLU-Pro

78.2%

Hard academic questions across 14 subjects — STEM, law, health, economics, philosophy and more. Professionally translated to German, with up to ten answer options per question.

10-option multiple choice translated · Professional translation

via li-lab/MMLU-ProX ↗

MMMLU

83.7%

OpenAI's multilingual MMLU, German split — general knowledge spanning STEM, the humanities, social sciences and other domains. Professionally translated to German.

4-option multiple choice translated · Professional translation

via openai/MMMLU ↗

MuSR

83.7%

Multi-step soft reasoning over long narrative contexts — murder mysteries, object placement and team allocation. Requires chaining clues across several paragraphs to reach the correct answer. Translated to German from the original English MuSR benchmark.

2–5 option multiple choice translated · Professional translation

via zayne-sprague/MuSR ↗

SB10K

15.2%

Native German social-media sentiment classification — positive, neutral or negative. Human-annotated German text, not translated. Run reasoning-off; scored as exact-match accuracy on the predicted label.

3-class sentiment native · Native German

via SB10K (via EuroEval) ↗

ScaLA

66.3%

Native German linguistic acceptability — does the sentence read as grammatical German (ja / nein)? Built from clean vs. minimally-corrupted German sentences. Run reasoning-off.

Binary acceptability native · Native German

via ScaLA-de (via EuroEval) ↗

Cost & speed

$0.174 per 1,000 questions

46 tokens / second

1.16s time to first token

Compare with other models → ← View leaderboard