遇见数据集

COMBINED DATASET

收藏
Zenodo2025-12-19 更新2026-05-26 收录
官方服务:

资源简介:

The dataset used in this study comprises 82,946 labeled news statements collected from multiple publicly available fact-checking and news verification sources focusing on Indian and Indic language content. Each instance is annotated for binary classification, with labels Real (0) and Fake (1). The dataset exhibits a natural class imbalance, consisting of 53,713 Fake samples (64.76%) and 29,233 Real samples (35.24%). The corpus spans a diverse set of languages, including English, Hindi, Tamil, Gujarati, Malayalam, Punjabi, Bengali, Telugu, Marathi, Nepali, and other low-resource languages. It also contains romanized and code-mixed text, reflecting realistic social media usage patterns in multilingual Indian settings. Language identifiers were retained to support language-wise evaluation. Data from different sources were merged into a unified format, retaining only semantically meaningful fields: news text, label, and language. The dataset’s scale, linguistic diversity, and presence of code-mixing make it suitable for evaluating multilingual transformer models for Indic fake news detection.

提供机构:
Zenodo
创建时间:
2025-12-19
二维码
社区交流群
二维码
科研交流群
商业服务