遇见数据集

MANITWEET

收藏
arXiv2023-05-24 更新2024-08-06 收录
数据链接:
官方服务:

资源简介:

MANITWEET数据集由伊利诺伊大学厄巴纳-香槟分校创建,包含3636条推文及其对应的新闻文章,旨在识别社交媒体上的新闻信息操纵。数据集内容丰富,涉及多种新闻主题,通过两轮人工标注确保数据质量。创建过程中,研究人员利用大型语言模型生成推文,并通过人工验证提高数据准确性。该数据集主要用于研究社交媒体上的信息操纵问题,特别是检测和分析新闻内容的篡改,以支持更准确的信息传播和事实核查。

The MANITWEET dataset, created by the University of Illinois Urbana-Champaign, consists of 3,636 pairs of tweets and their corresponding news articles, aiming to identify news information manipulation on social media. It covers diverse news topics and ensures data quality through two rounds of manual annotation. During its creation, researchers employed large language models to generate tweets and enhanced data accuracy via manual verification. This dataset is primarily utilized for researching information manipulation on social media, specifically for detecting and analyzing news content tampering, to support more accurate information dissemination and fact-checking.

创建时间:
2023-05-24
二维码
社区交流群
二维码
科研交流群
商业服务