遇见数据集

BIDWESH: A Bangla Regional Based Hate Speech Detection dataset

收藏
NIAID Data Ecosystem2026-05-02 收录
官方服务:

资源简介:

The BIDWESH dataset is the first benchmark corpus for hate speech detection in Bangla regional dialects, covering Noakhali, Chittagong, and Barishal. It consists of 9,183 manually translated and annotated instances derived from the BD-SHS dataset, ensuring balanced representation across dialects. Each sentence is labeled for hate or non-hate speech, with hate speech further annotated across 13 type categories and 7 target classes. The dataset preserves regional linguistic authenticity through native speaker translation and a rigorous five-stage validation process. BIDWESH supports multi-level classification, enabling nuanced analysis of hate expression in low-resource, dialectal Bangla contexts.

创建时间:
2025-07-21
二维码
社区交流群
二维码
科研交流群
商业服务