遇见数据集

Detecting East Asian Prejudice on Social Media

收藏
Zenodo2020-08-01 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

This repository contains: A deep learning model which distinguishes between Hostililty against East Asia, Criticism of East Asia, Discussion of East Asian prejudice and Neutral content. The F1 score is 0.83. A detailed annotation codebook used for marking up the tweets. A labelled dataset with 20,000 entries. A dataset with all 40,000 annotations, which can be used to investigate annotation processes for abusive content moderation. A list of thematic hashtag replacements. Three sets of annotations for the 1,000 most used hashtags in the original database of COVID-19 related tweets. Hashtags were annotated for COVID-19 relevance, East Asian relevance and stance. The outbreak of COVID-19 has transformed societies across the world as governments tackle the health, economic and social costs of the pandemic. It has also raised concerns about the spread of hateful language and prejudice online, especially hostility directed against East Asia. This data repository is for a classifier that detects and categorizes social media posts from Twitter into four classes: Hostility against East Asia, Criticism of East Asia, Meta-discussions of East Asian prejudice and a neutral class. The classifier achieves an F1 score of 0.83 across all four classes. We provide our final model (coded in Python), as well as a new 20,000 tweet training dataset used to make the classifier, two analyses of hashtags associated with East Asian prejudice and the annotation codebook. The classifier can be implemented by other researchers, assisting with both online content moderation processes and further research into the dynamics, prevalence and impact of East Asian prejudice online during this global pandemic. This work is a collaboration between The Alan Turing Institute and the Oxford Internet Institute. It was funded by the Criminal JusticeTheme of the Alan Turing Institute under Wave 1 of The UKRI Strategic Priorities Fund, EPSRC Grant EP/T001569/1

本数据集仓库包含以下内容:一款可区分针对东亚的敌意言论、对东亚的批评、对东亚偏见的讨论以及中立内容的深度学习模型,其F1分数为0.83;一份用于推文标注的详细标注编码手册(annotation codebook);一份包含20000条标注条目的标注数据集;一份包含全部40000条标注的数据集,可用于研究仇恨内容审核中的标注流程;一份主题话题标签(hashtag)替换列表;以及针对新冠疫情相关推文原始数据库中使用频率最高的1000个话题标签的三组标注,上述话题标签的标注维度包括与新冠疫情(COVID-19)的相关性、与东亚的相关性以及立场倾向。 新冠疫情(COVID-19)的爆发使全球社会发生深刻变革,各国政府需应对疫情带来的健康、经济与社会层面的多重代价,同时也引发了各界对网络仇恨言论与偏见传播的担忧,尤其是针对东亚的敌意言论。 本数据集仓库配套的分类器可检测并将Twitter(推特)上的社交媒体帖文划分为四大类别:针对东亚的敌意言论、对东亚的批评言论、关于东亚偏见的元讨论,以及中立类别。该分类器在全部四类任务上的F1分数可达0.83。 我们提供了最终的Python(Python)编写的模型,以及用于训练该分类器的全新20000条推文训练数据集、两份针对与东亚偏见相关话题标签的分析报告,以及前述标注编码手册。其他研究人员可部署该分类器,助力网络内容审核流程,以及针对本次全球疫情期间网络东亚偏见的动态、传播范围与影响展开进一步研究。 本研究由艾伦·图灵研究所(The Alan Turing Institute)与牛津互联网研究院(Oxford Internet Institute)合作完成,获艾伦·图灵研究所刑事司法主题项目资助,项目隶属于英国研究与创新署(UKRI)战略优先级基金第一批次,资助编号为EPSRC Grant EP/T001569/1。

提供机构:
Zenodo
创建时间:
2020-05-08
二维码
社区交流群
二维码
科研交流群
商业服务