遇见数据集

ORFND: A Source-Verified Odia Real–Fake News Dataset for Binary Fake-News Detection

收藏
Zenodo2026-08-13 更新2026-08-20 收录
官方服务:

资源简介:

ORFND: A Source-Verified Odia Real–Fake News Dataset for Binary Fake-News Detection ORFND is a source-verified dataset developed for research on binary fake-news detection in the Odia language. The dataset contains Odia news articles categorized into two classes: REAL and FAKE. It is intended to support research in Natural Language Processing (NLP), machine learning, misinformation detection, fact-checking, and low-resource Indian language processing. The dataset was constructed by collecting Odia news content from identifiable online news sources and compiling fake-news instances based on available fact-checking and verification evidence. Real-news instances were associated with identifiable news sources, while fake-news instances were included based on available verification or fact-checking evidence. The collection and verification process was designed to improve the reliability and traceability of the dataset. The dataset includes textual news information and associated metadata such as news title, article text, source, label, and relevant verification information where available. The binary labels represent REAL news and FAKE news. Before inclusion, the collected records were subjected to data cleaning and quality-control procedures, including removal of duplicate records, handling of incomplete records, normalization of textual content, and inspection of the Odia Unicode representation. Records with corrupted or unusable textual content were excluded from the final dataset. ORFND is intended primarily for academic and research purposes, including supervised fake-news classification, comparative evaluation of machine-learning and deep-learning models, Odia NLP research, misinformation detection, and explainable artificial intelligence. The dataset represents a research snapshot of Odia news and fact-checked information collected during the dataset construction period. Therefore, it may not represent all forms of misinformation or all Odia news sources. Researchers should consider temporal, source, and domain coverage limitations when interpreting experimental results. This dataset is released as a research resource to facilitate reproducible research in Odia fake-news detection and to encourage further development of reliable NLP resources for low-resource Indian languages.

提供机构:
Zenodo
创建时间:
2026-08-13
二维码
社区交流群
二维码
科研交流群
商业服务