unlearning-cleanslate/generations-10-llama-3_1-8b-simnpo-gentle-bm25-6t-target-100-checkpoint-187
收藏资源简介:
该数据集是一个多任务评估集合,主要用于评估语言模型在推理和少样本学习上的性能。它包含ARC挑战(AI2推理挑战)和BBH(BIG-Bench Hard)任务的变体,涉及布尔表达式、因果判断、日期理解、消歧问答、Dyck语言、形式谬误、几何形状、超常语序、逻辑演绎(三、五、七对象)、电影推荐、多步算术、导航、对象计数、企鹅表格、彩色对象推理和名字破坏等多种任务。每个任务配置包括问题输入、目标答案、生成参数(如采样设置)、模型响应、过滤后响应、评估指标和哈希值。数据集仅提供训练集分割,用于模型评估和基准测试。
This dataset is a multi-task evaluation collection designed primarily to assess the performance of language models on reasoning and few-shot learning. It includes variants of the ARC Challenge (AI2 Reasoning Challenge) and BBH (BIG-Bench Hard) tasks, covering areas such as boolean expressions, causal judgment, date understanding, disambiguation QA, Dyck languages, formal fallacies, geometric shapes, hyperbaton, logical deduction (three, five, and seven objects), movie recommendation, multistep arithmetic, navigation, object counting, penguins in a table, reasoning about colored objects, and ruin names. Each task configuration features input questions, target answers, generation parameters (e.g., sampling settings), model responses, filtered responses, evaluation metrics, and hash values. The dataset only provides a training split, intended for model evaluation and benchmarking.



