遇见数据集

Indic Media Bias Detection Dataset

收藏
Zenodo2026-01-15 更新2026-05-26 收录
官方服务:

资源简介:

Media bias classification remains an under-researched problem for low-resource languages. In this paper, we introduce the first exhaustive annotated dataset consisting of 300 unique articles from two leading Indian news agencies publishing in Hindi. We focus on Hindi, the most spoken and an official language of India. The study of media bias is particularly relevant in the Indian context due to the country’s linguistic diversity and large population, where the same event is often reported by multiple news agencies in different languages, thereby introducing scope for bias. The dataset comprises Hindi political news articles annotated by nine expert annotators and classified into nine bias categories when an article is identified as biased. For each biased article, annotators also provide a textual explanation, enabling future automated approaches to incorporate explainability when language models reason over the dataset. Cohen’s Kappa scores were 0.81 for simple bias/no-bias classification and 0.60 for the top-three average bias score, reflecting moderate to strong inter-annotator agreement.

提供机构:
Zenodo
创建时间:
2026-01-15
二维码
社区交流群
二维码
科研交流群
商业服务