German Language LLM Index
a PeerBench project

German LLM Benchmark · Model Profile

GPT-5.6 Luna

OpenAI provider-internal run 2026-08-13
75 .6%

avg. German score

+2.1pp above avg.

Benchmark breakdown

GermEval

81.9%

Native German named-entity recognition — identify persons, locations, organisations and misc entities in German text, emitted as JSON. Scored with seqeval micro-F1 excluding the noisy MISC class. Run reasoning-off.

Named-entity recognition native · Native German
via GermEval (via EuroEval) ↗ View run details →

INCLUDE

86.0%

Native German exam and licensing questions covering region-specific knowledge — history, law, civics and culture. Written by humans in German, not translated.

4-option multiple choice native · Native German
via CohereLabs/include-base-44 ↗ View run details →

MuSR

75.8%

Multi-step soft reasoning over long narrative contexts — murder mysteries, object placement and team allocation. Requires chaining clues across several paragraphs to reach the correct answer. Translated to German from the original English MuSR benchmark. Public results use the corrected v1.1 evaluation with 553 effective test cases.

2–5 option multiple choice translated · German translation
via TAUR-Lab/MuSR ↗ View run details →

SB10K

61.9%

Native German social-media sentiment classification — positive, neutral or negative. Human-annotated German text, not translated. Run reasoning-off; headline score is MCC on the predicted label.

3-class sentiment native · Native German
via SB10K (via EuroEval) ↗ View run details →

ScaLA

72.2%

Native German linguistic acceptability — does the sentence read as grammatical German (ja / nein)? Built from clean vs. minimally-corrupted German sentences. Run reasoning-off.

Binary acceptability native · Native German
via ScaLA-de (via EuroEval) ↗ View run details →

Cost & speed

$0.135 per 1,000 questions