遇见数据集

BanglaHateBench-12K: A Bangla/Banglish Hate Speech and Cyberbullying Detection Benchmark

收藏
Zenodo2026-06-08 更新2026-06-12 收录
官方服务:

资源简介:

BanglaHateBench-12K is a structured benchmark scaffold for hate speech and cyberbullying detection in Bangla, Banglish (romanized Bengali), and code-mixed text. The dataset contains 12,000 synthetic masked comments distributed across six categories: Normal, Offensive, Cyberbullying, Hate Speech, Threat, and Other Harassment/Abuse.The corpus covers three language varieties — Bangla (5,500), Banglish (5,000), and Mixed (1,500) — and is split into train (8,400), validation (1,800), and test (1,800) sets using a leakage-safe stratified strategy.This release (v0.1.0) is a silver scaffold: labels are algorithmically assigned and have not yet undergone full two-annotator human validation. It is intended for pipeline development, model prototyping, annotation workflow testing, and transformer fine-tuning experiments. Included resources cover the full annotation assignment structure, preprocessing pipelines, fast baseline models (SGD-logistic regression and Linear SVM with hashing vectorizer), and transformer training scripts compatible with BanglaBERT, mBERT, and XLM-R.A gold-validated version with verified inter-annotator agreement (Cohen's Kappa) is planned as a future release. Researchers using this dataset should clearly disclose its silver status in any published work.

提供机构:
Zenodo
创建时间:
2026-06-08
二维码
社区交流群
二维码
科研交流群
商业服务