BIG-Bench Hard (BBH) Data Quality Investigation Report
收藏资源简介:
38 ground truth labeling errors identified in the BIG-Bench Hard (BBH) benchmark dataset through systematic multi-model cross-verification. Verification Models GPT-4o (OpenAI) GPT-5.2 Pro (OpenAI, manual verification) Claude Sonnet 4 (Anthropic) Claude Opus 4.5 (Anthropic) DeepSeek-Reasoner (DeepSeek) Qwen3-8B (Alibaba) Error Summary date_understanding: 7 errors (calculation/date arithmetic) geometric_shapes: 31 errors (systematic K mislabeling) 17 cases: (K) ellipse → (A) circle (SVG arc with rx=ry) 14 cases: (K) trapezoid → (H) rectangle (perpendicular sides) Note on date_understanding GPT-5.2 Pro verified 8 disputed cases and confirmed the original labels were correct. The confirmed 7 errors are from unanimous multi-model agreement. Impact 12.4% error rate in geometric_shapes ground truth causes models with correct reasoning to be penalized.



