遇见数据集

KI-Benchmark-Deutsch

收藏
OpenML2026-07-20 更新2026-08-16 收录
官方服务:

资源简介:

Results of the recurring i6eal KI-Benchmark Deutsch (German Business AI Benchmark): leading large language models evaluated on real German business tasks (run 2026-06-22, 7 models, 6 categories, 24 tasks per model, scores 0-100). Each row is one model x category result. Columns: - run_date: date of the benchmark run (YYYY-MM-DD) - model_id: stable model identifier - model_name: human-readable model name (e.g. Claude Opus 4.8, GPT-5.5) - provider: model vendor (Anthropic, OpenAI, Google, ...) - category_id: benchmark category identifier (business, amtsdeutsch, recht, zusammenfassung, rag, sprache) - category_name_de / category_name_en: category names in German and English - score: category score 0-100 (higher is better) Categories cover formal business correspondence, official/bureaucratic German (Amtsdeutsch), verifiable German law and tax questions, faithful summarisation, source-grounded answering (RAG) and language quality. Subjective categories are scored by a cross-vendor panel of neutral judge models (name-blind, spot-checked by hand); reference categories are scored against ground truth - hallucinated statutes or deadlines count as wrong. Only sample prompts are public; the held-out test set stays private to prevent contamination. Key finding of this run: Claude Opus 4.8 leads overall with 98/100; official/bureaucratic German is the weakest category for every model family (category average 83.4). Canonical DOI (always latest version): 10.5281/zenodo.21452970. Live leaderboard and full methodology: https://i6eal.de/ki-benchmark-deutsch/ - Licence CC BY 4.0, attribution to i6eal.de.

创建时间:
2026-07-20
二维码
社区交流群
二维码
科研交流群
商业服务