BANTH
收藏资源简介:
BANTH数据集是由Penta Global Limited创建的,专门用于检测和分类转写孟加拉语中的仇恨言论。该数据集包含37,350个样本,来源于YouTube评论,每个样本都标记了多个目标群体,反映了区域人口统计特征。数据集的创建过程包括从YouTube API抓取评论、数据过滤、清洗和多轮注释验证。BANTH数据集的应用领域主要集中在仇恨言论的检测和分类,旨在解决低资源语言中仇恨言论自动检测的挑战。
The BANTH dataset, developed by Penta Global Limited, is specifically tailored for the detection and classification of hate speech in transliterated Bengali. It consists of 37,350 samples sourced from YouTube comments, with each sample annotated with multiple target groups that reflect regional demographic characteristics. The dataset creation workflow includes scraping comments via the YouTube API, data filtering, data cleaning, and multi-round annotation validation. Primarily applied in hate speech detection and classification, the BANTH dataset is designed to tackle the challenges of automatic hate speech detection in low-resource languages.

- 1BANTH: A Multi-label Hate Speech Detection Dataset for Transliterated BanglaPenta Global Limited · 2024年



