arpandeepk/generations-olmo-3-1025-7b-simnpo-gentle-checkpoint-64
收藏资源简介:
该数据集包含多个配置,用于评估AI模型的推理能力。主要配置包括:1) arc_challenge:基于AI2推理挑战的数据集,包含多项选择题,测试科学推理能力;2) bbh_cot_fewshot_*系列:基于Big-Bench Hard任务的思维链少样本评估数据集,涵盖布尔表达式、因果判断、日期理解、消歧问答、Dyck语言、形式谬误、几何形状、超序、逻辑演绎(三、五、七对象)、电影推荐、多步算术、导航、对象计数、表格中的企鹅、彩色对象推理和名称毁坏等多种推理任务。数据集特征包括文档ID、输入问题、目标答案、生成参数、模型响应、过滤响应、评估指标和分数等,用于全面评估模型性能。
This dataset includes multiple configurations for evaluating AI model reasoning capabilities. Key configurations are: 1) arc_challenge: A dataset based on the AI2 Reasoning Challenge, containing multiple-choice questions to test scientific reasoning; 2) bbh_cot_fewshot_* series: Chain-of-thought few-shot evaluation datasets for Big-Bench Hard tasks, covering reasoning domains such as boolean expressions, causal judgement, date understanding, disambiguation QA, Dyck languages, formal fallacies, geometric shapes, hyperbaton, logical deduction (three, five, seven objects), movie recommendation, multistep arithmetic, navigation, object counting, penguins in a table, reasoning about colored objects, and ruin names. Dataset features include document ID, input questions, target answers, generation arguments, model responses, filtered responses, evaluation metrics, and scores, providing comprehensive performance assessment.




