遇见数据集

Multilingual Counterfactual Dataset

收藏
arXiv2021-09-16 更新2024-06-21 收录
官方服务:

资源简介:

Multilingual Counterfactual Dataset是由亚马逊和利物浦大学合作创建的多语言数据集,专注于产品评论中的反事实陈述。该数据集涵盖英语、德语和日语,包含10000条评论,旨在通过高质量的专业标注提升反事实检测模型的准确性。数据集的创建过程包括从亚马逊客户评论数据集中筛选评论,并通过两轮迭代选择候选句子进行人工标注。该数据集的应用领域主要集中在自然语言处理,特别是信息检索和情感分析中,以帮助识别和处理包含反事实信息的文本。

The Multilingual Counterfactual Dataset is a multilingual dataset co-created by Amazon and the University of Liverpool, focusing on counterfactual statements in product reviews. Covering English, German and Japanese, the dataset contains 10,000 reviews and aims to improve the accuracy of counterfactual detection models through high-quality professional annotations. The dataset's creation process includes screening reviews from the Amazon Customer Reviews Dataset, and selecting candidate sentences for manual annotation via two rounds of iteration. Its application areas are mainly focused on natural language processing, particularly information retrieval and sentiment analysis, to assist in identifying and processing texts containing counterfactual information.

创建时间:
2021-04-14
搜集汇总
数据集介绍
Multilingual Counterfactual Dataset 数据集图片
背景与挑战
背景概述
该数据集是一个用于反事实检测(CFD)二元分类任务的多语言数据集,包含来自亚马逊产品评论的英语、德语和日语句子,标注为反事实语句(描述未发生或不可能发生的事件)。其特点包括由专业语言学家进行高质量标注,并提供标注指南和线索词列表以支持研究。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务