遇见数据集

CRASS

收藏
arXiv2022-10-05 更新2024-06-21 收录
官方服务:

资源简介:

CRASS数据集是由莱比锡大学应用信息学研究所创建,用于评估大型语言模型在反事实推理能力上的表现。该数据集包含274个高质量的反事实条件语句(PCT),通过亚马逊Mechanical Turk平台生成和验证。数据集设计用于测试模型对假设情景的理解和推理,特别是在处理反事实条件时。CRASS数据集的应用领域主要集中在提升语言模型在复杂逻辑推理任务上的性能,特别是在理解和生成反事实条件语句方面。

The CRASS dataset was developed by the Institute of Applied Informatics at Leipzig University to evaluate the counterfactual reasoning capabilities of large language models (LLMs). It contains 274 high-quality counterfactual conditional statements (PCTs), which were generated and validated via the Amazon Mechanical Turk platform. This dataset is designed to test models' understanding and reasoning regarding hypothetical scenarios, especially when dealing with counterfactual conditions. The primary application scope of the CRASS dataset focuses on improving the performance of language models in complex logical reasoning tasks, particularly in understanding and generating counterfactual conditional statements.

创建时间:
2021-12-22
搜集汇总
数据集介绍
CRASS 数据集图片
背景与挑战
背景概述
CRASS是一个用于评估大型语言模型反事实推理能力的数据集和基准测试。它包含前提-反事实元组(PCTs),通过对比反事实条件与基础前提来设计任务,支持固定目标模式和开放评分模式两种评估方式。该数据集旨在为模型性能提供新的测试工具,并已集成到BIG Bench中,相关研究已在LREC 2022发表。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务