遇见数据集

HS-BAN

收藏
arXiv2021-12-03 更新2024-08-06 收录
数据链接:
官方服务:

资源简介:

HS-BAN是由沙贾拉尔科技大学创建的孟加拉语仇恨言论数据集,包含超过50,000条来自Facebook和YouTube的评论,其中40.17%为仇恨言论。数据集通过严格的标注指南减少人为标注偏差,并进行了语言学预处理以提取不同类型的俚语。该数据集旨在解决社交媒体上仇恨言论检测的问题,特别是在孟加拉语环境中,通过提供一个大规模、语言多样化的数据集来支持相关研究。

HS-BAN is a Bengali hate speech dataset developed by Shahjalal University of Science and Technology. It contains over 50,000 comments collected from Facebook and YouTube, with 40.17% of them classified as hate speech. To mitigate human annotation bias, the dataset follows strict annotation guidelines and has been subjected to linguistic preprocessing to extract various types of slang. This dataset is designed to tackle the issue of hate speech detection on social media, particularly in the Bengali language context, by offering a large-scale, linguistically diverse dataset to support related research.

创建时间:
2021-12-03
二维码
社区交流群
二维码
科研交流群
商业服务