遇见数据集

FactCheckMeta-SP: A Metadata-Only Fact-Checking Dataset from Snopes and PolitiFact for Automated Fact-Checking Research

收藏
Zenodo2026-06-25 更新2026-06-28 收录
官方服务:

资源简介:

Dataset Description These datasets are structured collections of fact-checking records gathered from publicly accessible fact-checking platforms. They have been designed to support research in automated fact-checking, misinformation detection, information retrieval, natural language processing (NLP), computational journalism, and machine learning. The public release contains metadata only and does not include the full text of the original fact-checking articles. Data range for news in Snopes2025 from 2025-07-14 to 2025-11-30. Data range for news in Politifact2025 from 2025-01-02 to 2025-11-26. Data Format The dataset is distributed in JSON format, where each record represents a single fact-checking instance. Example: { "label": "real", "news_id": "politifact_9187", "news_source": "Politifact", "title": "In Congress, did Sean Duffy vote against upgrading air traffic control systems? Here’s what we know", "url": "https://www.politifact.com/factchecks/2025/may/21/robert-reich/Sean-Duffy-air-traffic-control-congress/", "published_date": "2025-05-21" } Schema Field Type Description label string Veracity label assigned by the source fact-checking organization. news_id string Unique identifier generated for the dataset. news_source string Name of the originating fact-checking platform. title string Title of the fact-checking article. url string URL of the original fact-checking article. published_date string Publishing date of the news. Data Collection Records were collected from publicly accessible fact-checking websites and normalized into a unified schema. Source-specific verdicts were mapped into a common labeling framework to facilitate cross-source analysis and machine learning applications. The dataset is intended for: Automated fact-checking research Misinformation detection Retrieval-Augmented Generation (RAG) evaluation Claim verification Information retrieval experiments Benchmark construction Computational journalism studies Dataset Characteristics Metadata-only distribution Unified schema across multiple fact-checking sources Source attribution preserved through URLs and source identifiers Suitable for large-scale data analysis and benchmarking Compatible with Python, Pandas, Hugging Face Datasets, and common machine learning pipelines Limitations This release does not contain the original article content, images, or other copyrighted materials. Researchers requiring access to the full fact-checking articles should retrieve them directly from the original sources using the provided URLs. Veracity labels are inherited from the original fact-checking organizations and may reflect source-specific editorial methodologies. Users should consult the original articles for complete contextual information. This dataset has been developed to support research on automated fact-checking, misinformation detection, and the evaluation of natural language processing (NLP) and artificial intelligence (AI) systems. The dataset contains structured metadata collected from publicly accessible fact-checking websites. Each record includes the verdict label assigned by the original fact-checking organization, a unique identifier, the source organization, the title of the fact-check article, and a URL pointing to the original source. To respect intellectual property rights and promote responsible data sharing, this public release does not include the full text, images, or other copyrighted content from the original fact-checking articles. Researchers interested in the original content should access it directly through the URLs provided in the dataset. The dataset was created within an academic research context and is intended exclusively for scientific research, educational activities, and the evaluation of fact-checking and misinformation detection systems. The data collection process was conducted in reliance on the text and data mining (TDM) exception provided by applicable copyright laws implementing Article 3 of Directive (EU) 2019/790 on Copyright in the Digital Single Market (DSM Directive), or equivalent provisions where applicable. All intellectual property rights in the original source materials remain with their respective owners. Inclusion of metadata, labels, titles, and source URLs in this dataset does not imply any transfer of ownership or licensing of the underlying works. Rights Holder Notice and Takedown Procedure If any rights holder believes that material contained in this dataset infringes their intellectual property rights or other legal rights, they may contact the dataset curators and request its removal. Upon receiving a substantiated request, the curators will review the claim and, where appropriate, remove or modify the relevant material within a reasonable timeframe.

提供机构:
Zenodo
创建时间:
2026-06-25
二维码
社区交流群
二维码
科研交流群
商业服务