Indic Media Bias Detection Dataset
收藏资源简介:
Media bias classification remains an under-researched problem for low-resource languages. In this paper, we introduce the first exhaustive annotated dataset consisting of 300 unique articles from two leading Indian news agencies publishing in Hindi. We focus on Hindi, the most spoken and an official language of India. The study of media bias is particularly relevant in the Indian context due to the country’s linguistic diversity and large population, where the same event is often reported by multiple news agencies in different languages, thereby introducing scope for bias. The dataset comprises Hindi political news articles annotated by nine expert annotators and classified into nine bias categories when an article is identified as biased. For each biased article, annotators also provide a textual explanation, enabling future automated approaches to incorporate explainability when language models reason over the dataset. Cohen’s Kappa scores were 0.81 for simple bias/no-bias classification and 0.60 for the top-three average bias score, reflecting moderate to strong inter-annotator agreement.



