COMMUNITYNOTES
收藏资源简介:
COMMUNITYNOTES是一个大规模的多语言数据集,包含104,966条可能具有误导性的帖子及其对应的用户提供的解释性笔记和有用性标签。该数据集由墨尔本大学和MBZUAI的研究团队创建,旨在探索社区注释中解释性笔记的有用性及其原因。数据集内容涵盖了英语和其他多种语言,其中英语帖子占主导地位。该数据集通过从X社区注释网站上收集数据,并使用官方注释算法计算每个笔记的综合有用性和原因标签而构建。数据集已被分割为训练集、开发集和测试集,用于评估预测笔记有用性和原因的任务。该数据集的应用领域为社区事实核查,旨在解决如何提高社区注释的有用性和可解释性问题。
COMMUNITYNOTES is a large-scale multilingual dataset comprising 104,966 potentially misleading posts, along with their corresponding user-provided explanatory notes and usefulness labels. Developed by a research team from the University of Melbourne and MBZUAI, this dataset aims to explore the usefulness of explanatory notes in community annotations and the underlying reasons for such usefulness. The dataset covers English and multiple other languages, with English posts constituting the majority. It is built via data collection from the X Community Notes website and the computation of comprehensive usefulness and reason labels for each note using the official annotation algorithm. The dataset has been partitioned into training, development, and test sets to support the evaluation of tasks focused on predicting note usefulness and its corresponding reasons. Its application domain lies in community fact-checking, with the core goal of addressing the challenge of improving the usefulness and interpretability of community annotations.




