BANTH
收藏资源简介:
BANTH数据集是由Penta Global Limited和Islamic University of Technology合作创建的,专门用于检测和分类转写孟加拉语中的仇恨言论。该数据集包含37,350条样本,主要来源于YouTube评论,涵盖新闻与政治、人物与博客、娱乐等多个类别。数据集的创建过程包括数据抓取、过滤、清洗和多轮人工标注与验证,确保了数据的高质量和准确性。BANTH数据集的应用领域主要集中在多标签仇恨言论检测,旨在解决低资源语言中仇恨言论自动检测的挑战,并为未来的跨语言和多标签分类研究奠定基础。
The BANTH dataset was co-created by Penta Global Limited and Islamic University of Technology, specifically designed for detecting and classifying hate speech in transcribed Bengali. This dataset contains 37,350 samples, mainly sourced from YouTube comments, covering multiple categories including news and politics, figures and blogs, entertainment, etc. The creation process of the BANTH dataset includes data crawling, filtering, cleaning, as well as multiple rounds of manual annotation and verification, ensuring the high quality and accuracy of the data. The application fields of the BANTH dataset mainly focus on multi-label hate speech detection, aiming to address the challenges of automatic hate speech detection in low-resource languages, and lay a foundation for future cross-lingual and multi-label classification research.




