unlearning-cleanslate/generations-nemotron-nano-9b-v2-simnpo-baseline
收藏资源简介:
该数据集包含多个配置,用于评估和测试AI模型的推理与认知能力。主要配置包括:1. arc_challenge:AI推理挑战数据集,包含问题、答案选项和正确答案,用于测试模型的多选题推理能力,有1172个训练样本。2. bbh_cot_fewshot_*系列:基于Big-Bench Hard任务的思维链少样本数据集,涵盖多种认知任务,如布尔表达式、因果判断、日期理解、消歧问答、Dyck语言、形式谬误、几何形状、超常语序、逻辑演绎(三/五/七个对象)、电影推荐、多步算术、导航、对象计数、表格中的企鹅、有色物体推理和名字破坏等,每个任务通常有250个训练样本(部分有差异)。数据集特征包括输入文本、目标输出、生成参数、模型响应、过滤响应、哈希值和评分等,旨在支持模型在少样本设置下的思维链推理评估。
This dataset comprises multiple configurations for evaluating and testing the reasoning and cognitive capabilities of AI models. The core configurations are as follows: 1. arc_challenge: An AI reasoning challenge dataset containing questions, answer options, and correct answers, designed to test the multiple-choice reasoning ability of models, with 1172 training samples. 2. bbh_cot_fewshot_* series: A chain-of-thought few-shot dataset based on Big-Bench Hard tasks, covering a diverse set of cognitive tasks including boolean expressions, causal judgment, date comprehension, disambiguated question answering, Dyck languages, formal fallacies, geometric shapes, non-standard word order, logical deduction (involving three, five, or seven objects), movie recommendation, multi-step arithmetic, navigation, object counting, penguins in tables, colored object reasoning, and name scrambling. Each task typically includes 250 training samples, with individual variations for some tasks. The dataset features include input text, target output, generation parameters, model responses, filtered responses, hash values, and scores, etc., which is intended to support the evaluation of chain-of-thought reasoning for models under few-shot settings.




