BanglaHateBench-12K: A Bangla/Banglish Hate Speech and Cyberbullying Detection Benchmark
收藏资源简介:
BanglaHateBench-12K is a structured benchmark scaffold for hate speech and cyberbullying detection in Bangla, Banglish (romanized Bengali), and code-mixed text. The dataset contains 12,000 synthetic masked comments distributed across six categories: Normal, Offensive, Cyberbullying, Hate Speech, Threat, and Other Harassment/Abuse.The corpus covers three language varieties — Bangla (5,500), Banglish (5,000), and Mixed (1,500) — and is split into train (8,400), validation (1,800), and test (1,800) sets using a leakage-safe stratified strategy.This release (v0.1.0) is a silver scaffold: labels are algorithmically assigned and have not yet undergone full two-annotator human validation. It is intended for pipeline development, model prototyping, annotation workflow testing, and transformer fine-tuning experiments. Included resources cover the full annotation assignment structure, preprocessing pipelines, fast baseline models (SGD-logistic regression and Linear SVM with hashing vectorizer), and transformer training scripts compatible with BanglaBERT, mBERT, and XLM-R.A gold-validated version with verified inter-annotator agreement (Cohen's Kappa) is planned as a future release. Researchers using this dataset should clearly disclose its silver status in any published work.



