Individual experiment result
GermEval NER (DE)
Claude Opus 4.8 Anthropic
Micro-F1 excluding MISC reasoning off provider-internal
85.8%
Micro-F1 excluding MISC
Run details
- Date
- 2026-06-13
- Test cases
- 1,024
- Median latency
- 2.12s
- Total cost
- $11.88
- Avg. prompt tokens
- 2,110.5
- Avg. answer tokens
- 42
- Avg. reasoning tokens
- —
- Quantization
- provider-internal
About this benchmark
Native German named-entity recognition (person/location/org/misc) emitted as JSON; seqeval micro-F1 excluding MISC.
Provenance
- Canonical model
- anthropic/claude-opus-4.8
- Run ID
2026-06-13T11-17-44__openai__cc__claude-opus-4-8__germeval_de