reddit-br-toxicity-dataset
收藏资源简介:
该数据集包含2,500个从巴西最受欢迎的Reddit社区中手动标注的评论,用于在线社交网络中的毒性内容检测。数据集通过众包方式由计算机科学系和UFMG的语言学团队共同标注。数据集旨在促进针对低资源语言如巴西葡萄牙语的机器学习技术的实验和进步。
This dataset comprises 2,500 manually annotated comments from the most popular Reddit communities in Brazil, designed for the detection of toxic content in online social networks. The annotations were collaboratively conducted through crowdsourcing by the Department of Computer Science and the linguistics team at UFMG. The dataset aims to foster experimentation and advancement in machine learning techniques for low-resource languages such as Brazilian Portuguese.
数据集概述
数据集名称
Toxic Content Detection in online social networks: a new dataset from Brazilian Reddit Communities
数据集内容
- 样本数量: 2,500条手动标注的评论
- 数据来源: 从巴西Reddit社区中最大的10个子论坛提取
- 数据收集时间: 2022年1月至2022年12月
- 数据类型: 评论文本及其毒性标签
数据集结构
- 字段:
id: 评论在Reddit平台的唯一标识body: 原始评论文本is_toxic: 评论的最终标签(0表示非毒性,1表示毒性,-1表示标注者意见不一致)
标注过程
- 标注者: 来自计算机科学系(DCC)和UFMG的语言学组
- 标注方法: 通过众包方式进行,标注者将评论标记为
Toxic,Non-toxic,I do not know,Missing info - 最终标签: 通过多数投票确定
数据可用性
- 文件格式: CSV
- 文件路径:
dataset/toxicity_br_labeled_data.csv
数据集用途
- 旨在促进针对低资源语言(如巴西葡萄牙语)的毒性分类模型的实验和进步,改进现有方法或提出新方法。
引用格式
cite Lima, Q. Luiz Henrique; Pagano, S. Adriana; da Silva, A.P.C. 2024. Toxic Content Detection in online social networks: a new dataset from Brazilian Reddit Communities. 16th International Conference on Computational Processing of Portuguese (PROPOR 2024).




