Writer/IRT-mislabeled-items
收藏资源简介:
该数据集通过项目反应理论(IRT)检测出潜在错误标记的基准项目,源自论文《使用项目反应理论审计LLM基准》。数据集中包含1001个训练样本,每个样本具有多个特征列,如唯一标识符、基准家族、子集名称、任务类型(偏好或多选)、IRT标志、弱参考标签标志、IRT分数差异、缺失原因、GPT-5.4的弱参考标签及其解释、参考答案和领先非一致答案的投票计数等。数据集来源于多个公开基准,包括用于偏好评估的RewardBench、RewardBench 2、RM-Bench和JudgeBench,以及用于事实性多项选择的GPQA Diamond、MATH和GSM8K。这些项目被标记的条件是IRT的delta_li大于0或GPT-5.4弱参考标签为mislabel或unsure,旨在帮助研究人员识别和审核基准数据中的标签问题。
This dataset is a collection of benchmark items with potential label errors detected via Item Response Theory (IRT), sourced from the paper *Using Item Response Theory to Audit LLM Benchmarks*. It contains 1,001 training samples, each with multiple feature columns including unique identifier, benchmark family, subset name, task type (preference or multiple-choice), IRT flag, weak reference label flag, IRT score difference, reason for missing values, weak reference label and its explanation from GPT-5.4, reference answer, and vote count of the leading non-consistent answer. The dataset is compiled from multiple public benchmarks: RewardBench, RewardBench 2, RM-Bench and JudgeBench for preference evaluation tasks, as well as GPQA Diamond, MATH and GSM8K for factual multiple-choice tasks. These items are flagged when either the IRT-derived delta_li value is greater than 0, or the weak reference label from GPT-5.4 is marked as "mislabel" or "unsure". This dataset aims to assist researchers in identifying and auditing label issues within benchmark datasets.




