NELA-GT-2019
收藏资源简介:
NELA-GT-2019是由伦斯勒理工学院创建的大型多标签新闻数据集,旨在研究新闻文章中的虚假信息。该数据集包含2019年1月1日至12月31日期间从260个来源收集的112万篇新闻文章。数据集内容丰富,涵盖主流和替代新闻来源,每篇文章都附有来自7个不同评估网站的源级真实性标签。创建过程中,研究人员通过定期抓取RSS feeds来收集数据,并从多个评估网站获取真实性标签。该数据集适用于新闻真实性研究,特别是机器学习和计算社会科学领域,有助于理解和检测新闻中的虚假信息。
NELA-GT-2019 is a large-scale multi-label news dataset created by Rensselaer Polytechnic Institute for research on disinformation in news articles. It encompasses 1.12 million news articles collected from 260 distinct sources between January 1 and December 31, 2019. The dataset features rich content, spanning both mainstream and alternative news outlets, with each article paired with source-level authenticity labels obtained from 7 different evaluation websites. During its development, researchers collected data by regularly crawling RSS feeds and acquired the authenticity labels from multiple evaluation platforms. This dataset is applicable to news authenticity-related research, especially in the fields of machine learning and computational social science, and helps advance the understanding and detection of disinformation in news articles.




