BIG-Bench Hard (BBH) Data Quality Investigation Report
收藏官方服务:
资源简介:
49 ground truth labeling errors identified in the BIG-Bench Hard (BBH) benchmark dataset through systematic multi-model cross-verification. Error Summary date_understanding: 7 errors geometric_shapes: 42 errors 22 cases: (K) ellipse → (A) circle 20 cases: (K) trapezoid → (H) rectangle Verification 4-phase verification using GPT-4o, GPT-5.2 Pro, Claude Sonnet/Opus, DeepSeek-Chat/Reasoner, and Qwen3-8B. Impact 16.8% error rate in geometric_shapes ground truth.
提供机构:
Zenodo创建时间:
2025-12-30



