CounterBench
收藏资源简介:
CounterBench是一个专为评估大型语言模型在反事实推理任务上的性能而设计的综合数据集。该数据集由Rutgers University创建,包含1000个反事实推理问题,涵盖不同的难度级别、因果图结构、反事实问题类型和非 sensical名称变体。数据集中的问题旨在通过要求真正的推理而不仅仅是模式识别或记忆响应,系统地评估四个关键维度。该数据集适用于医疗保健、商业、公共管理等领域,支持对错过的机会和替代结果进行评估,从而指导决策制定。
CounterBench is a comprehensive dataset specifically engineered to evaluate the performance of large language models (LLMs) on counterfactual reasoning tasks. Created by Rutgers University, it comprises 1000 counterfactual reasoning problems covering diverse difficulty levels, causal graph structures, counterfactual question types, and nonsensical name variants. The problems within this dataset are designed to systematically evaluate four key dimensions by requiring genuine reasoning rather than mere pattern recognition or memorized responses. This dataset is applicable across domains including healthcare, business, and public administration, supporting the assessment of missed opportunities and alternative outcomes to guide decision-making.




