unlearning-cleanslate/generations-20-DEBUG-llama-3_1-8b-simnpo-gentle-igm-10b-target-100-localtrain-checkpoint-1
收藏资源简介:
该数据集包含多个配置,用于评估语言模型的推理和问题解决能力。配置包括arc_challenge(AI推理挑战)、bbh_cot_fewshot_boolean_expressions(布尔表达式少样本思维链)、bbh_cot_fewshot_causal_judgement(因果判断少样本思维链)等,覆盖逻辑推理、数学、日期理解、对象计数、电影推荐等多种任务。每个示例包含输入问题、目标答案、生成参数、模型响应和评分,旨在测试模型在少样本或零样本设置下的表现。数据集结构详细,包括训练分割,示例数量从146到1172不等,适用于NLP研究和基准测试。
This dataset includes multiple configurations for evaluating language models reasoning and problem-solving abilities. Configurations encompass arc_challenge (AI reasoning challenge), bbh_cot_fewshot_boolean_expressions (Boolean expressions few-shot chain-of-thought), bbh_cot_fewshot_causal_judgement (causal judgement few-shot chain-of-thought), among others, covering tasks such as logical reasoning, mathematics, date understanding, object counting, movie recommendation, and more. Each example contains input questions, target answers, generation arguments, model responses, and scores, designed to test model performance in few-shot or zero-shot settings. The dataset features detailed structures, including training splits with example counts ranging from 146 to 1172, suitable for NLP research and benchmarking.




