Navigating News Narratives: A Media Bias Analysis Dataset
收藏资源简介:
The prevalence of bias in the news media has become a critical issue, affecting public perception on a range of important topics such as political views, health, insurance, resource distributions, religion, race, age, gender, occupation, and climate change. The media has a moral responsibility to ensure accurate information dissemination and to increase awareness about important issues and the potential risks associated with them. This highlights the need for a solution that can help mitigate against the spread of false or misleading information and restore public trust in the media. Data description: This is a dataset for news media bias covering different dimensions of the biases: political, hate speech, political, toxicity, sexism, ageism, gender identity, gender discrimination, race/ethnicity, climate change, occupation, spirituality, which makes it a unique contribution. The dataset used for this project does not contain any personally identifiable information (PII). Data Format: The format of data is: ID: Numeric unique identifier. Text: Main content. Dimension: Categorical descriptor of the text. Biased_Words: List of words considered biased. Aspect: Specific topic within the text. Label: Bias True/False value Aggregate Label: Calculated through multiple weighted formulae Annotation Scheme: The annotation scheme is based on Active learning, which is Manual Labeling --> Semi-Supervised Learning --> Human Verifications (iterative process) Bias Label: Indicate the presence/absence of bias (e.g., no bias, mild, strong). Words/Phrases Level Biases: Identify specific biased words/phrases. Subjective Bias (Aspect): Capture biases related to content aspects. List of datasets used : We curated different news categories like Climate crisis news summaries , occupational, spiritual/faith/ general using RSS to capture different dimensions of the news media biases. The annotation is performed using active learning to label the sentence (either neural/ slightly biased/ highly biased) and to pick biased words from the news. We also utilize publicly available data from the following links. Our Attribution to others. MBIC (media bias): Spinde, Timo, Lada Rudnitckaia, Kanishka Sinha, Felix Hamborg, Bela Gipp, and Karsten Donnay. "MBIC--A Media Bias Annotation Dataset Including Annotator Characteristics." arXiv preprint arXiv:2105.11910 (2021). https://zenodo.org/records/4474336 Hyperpartisan news: Kiesel, Johannes, Maria Mestre, Rishabh Shukla, Emmanuel Vincent, Payam Adineh, David Corney, Benno Stein, and Martin Potthast. "Semeval-2019 task 4: Hyperpartisan news detection." In Proceedings of the 13th International Workshop on Semantic Evaluation, pp. 829-839. 2019. https://huggingface.co/datasets/hyperpartisan_news_detection Toxic comment classification: Adams, C.J., Jeffrey Sorensen, Julia Elliott, Lucas Dixon, Mark McDonald, Nithum, and Will Cukierski. 2017. "Toxic Comment Classification Challenge." Kaggle. https://kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge. Jigsaw Unintended Bias: Adams, C.J., Daniel Borkan, Inversion, Jeffrey Sorensen, Lucas Dixon, Lucy Vasserman, and Nithum. 2019. "Jigsaw Unintended Bias in Toxicity Classification." Kaggle. https://kaggle.com/competitions/jigsaw-unintended-bias-in-toxicity-classification. Age Bias : Díaz, Mark, Isaac Johnson, Amanda Lazar, Anne Marie Piper, and Darren Gergle. "Addressing age-related bias in sentiment analysis." In Proceedings of the 2018 chi conference on human factors in computing systems, pp. 1-14. 2018. Age Bias Training and Testing Data - Age Bias and Sentiment Analysis Dataverse (harvard.edu) Multi-dimensional news Ukraine: Färber, Michael, Victoria Burkard, Adam Jatowt, and Sora Lim. "A multidimensional dataset based on crowdsourcing for analyzing and detecting news bias." In Proceedings of the 29th ACM International Conference on Information & Knowledge Management, pp. 3007-3014. 2020. https://zenodo.org/records/3885351#.ZF0KoxHMLtV Social biases: Sap, Maarten, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, and Yejin Choi. "Social bias frames: Reasoning about social and power implications of language." arXiv preprint arXiv:1911.03891 (2019). https://maartensap.com/social-bias-frames/ Goal of this dataset :We want to offer open and free access to dataset, ensuring a wide reach to researchers and AI practitioners across the world. The dataset should be user-friendly to use and uploading and accessing data should be straightforward, to facilitate usage. If you use this dataset, please cite us. Navigating News Narratives: A Media Bias Analysis Dataset © 2023 by Shaina Raza, Vector Institute is licensed under CC BY-NC 4.0
新闻媒体中的偏见问题已成为至关重要的公共议题,其对政治观点、健康、保险、资源分配、宗教、种族、年龄、性别、职业及气候变化等诸多重要领域的公众认知均造成影响。媒体肩负道义责任,需保障信息的准确传播,并提升公众对重要议题及其相关潜在风险的认知水平。这凸显出亟需一套解决方案,以助力遏制虚假或误导性信息的扩散,重建公众对媒体的信任。 ### 数据描述 本数据集面向新闻媒体偏见问题,涵盖多维度偏见范畴:政治偏见(political)、仇恨言论(hate speech)、政治偏见(political)、恶意言论(toxicity)、性别歧视(sexism)、年龄偏见(ageism)、性别认同(gender identity)、性别歧视(gender discrimination)、种族/民族偏见(race/ethnicity)、气候变化偏见(climate change)、职业偏见(occupation)、精神信仰偏见(spirituality),这使其具备独特的研究价值。本项目使用的数据集未包含任何个人可识别信息(PII, personally identifiable information)。 ### 数据格式 数据格式如下: - **ID**:唯一数值标识符 - **Text**:文本主体内容 - **Dimension**:文本的分类描述项 - **Biased_Words**:被认定为存在偏见的词汇列表 - **Aspect**:文本所涉及的具体主题 - **Label**:偏见存在与否的布尔值(True/False,即是否存在偏见) - **Aggregate Label**:通过多种加权公式计算得到的聚合标签 - **Annotation Scheme**:标注流程基于主动学习(Active learning),具体为「手动标注 → 半监督学习 → 人工核验」的迭代流程 - **Bias Label**:用于标识偏见的存在等级,例如无偏见、轻度偏见、重度偏见 - **Words/Phrases Level Biases**:用于识别特定的偏见性词汇或短语 - **Subjective Bias (Aspect)**:捕捉与内容主题相关的偏见 ### 数据集来源 我们通过RSS源采集了多类新闻数据,包括气候危机新闻摘要、职业类、精神/信仰/通用类新闻,以覆盖新闻媒体偏见的多维度特征。标注流程采用主动学习,对句子进行偏见等级标注(无偏见、轻度偏见、重度偏见),并从新闻文本中提取存在偏见的词汇。 此外,我们还使用了以下公开可得数据集,并对原作者致谢: 1. **MBIC(媒体偏见数据集)**:Spinde, Timo, Lada Rudnitckaia, Kanishka Sinha, Felix Hamborg, Bela Gipp, Karsten Donnay. "MBIC--A Media Bias Annotation Dataset Including Annotator Characteristics." arXiv预印本arXiv:2105.11910 (2021). 链接:https://zenodo.org/records/4474336 2. **Hyperpartisan新闻数据集**:Kiesel, Johannes, Maria Mestre, Rishabh Shukla, Emmanuel Vincent, Payam Adineh, David Corney, Benno Stein, Martin Potthast. "Semeval-2019 task 4: Hyperpartisan news detection." 发表于《第13届语义评估国际研讨会论文集》,第829-839页,2019年。链接:https://huggingface.co/datasets/hyperpartisan_news_detection 3. **恶意评论分类挑战赛数据集**:Adams, C.J., Jeffrey Sorensen, Julia Elliott, Lucas Dixon, Mark McDonald, Nithum, Will Cukierski. "Toxic Comment Classification Challenge." Kaggle, 2017年。链接:https://kaggle.com/competitions/jigsaw-toxic-comment-classification-challenge 4. **Jigsaw恶意偏见数据集**:Adams, C.J., Daniel Borkan, Inversion, Jeffrey Sorensen, Lucas Dixon, Lucy Vasserman, Nithum. "Jigsaw Unintended Bias in Toxicity Classification." Kaggle, 2019年。链接:https://kaggle.com/competitions/jigsaw-unintended-bias-in-toxicity-classification 5. **年龄偏见数据集**:Díaz, Mark, Isaac Johnson, Amanda Lazar, Anne Marie Piper, Darren Gergle. "Addressing age-related bias in sentiment analysis." 发表于《2018年CHI人机交互系统会议论文集》,第1-14页,2018年。数据来源:Age Bias Training and Testing Data - Age Bias and Sentiment Analysis Dataverse (harvard.edu) 6. **多维乌克兰新闻数据集**:Färber, Michael, Victoria Burkard, Adam Jatowt, Sora Lim. "A multidimensional dataset based on crowdsourcing for analyzing and detecting news bias." 发表于《第29届ACM信息与知识管理国际会议论文集》,第3007-3014页,2020年。链接:https://zenodo.org/records/3885351#.ZF0KoxHMLtV 7. **社会偏见数据集**:Sap, Maarten, Saadia Gabriel, Lianhui Qin, Dan Jurafsky, Noah A. Smith, Yejin Choi. "Social bias frames: Reasoning about social and power implications of language." arXiv预印本arXiv:1911.03891 (2019). 链接:https://maartensap.com/social-bias-frames/ ### 数据集目标 本数据集旨在开放免费获取,确保全球范围内的研究人员与人工智能从业者均可广泛使用。本数据集易用性强,数据上传与获取流程简便,以推动相关研究工作的开展。 若您使用本数据集,请引用以下文献: 《Navigating News Narratives: A Media Bias Analysis Dataset》© 2023 作者为Shaina Raza、Vector Institute,采用CC BY-NC 4.0许可协议。



