RoBangHate: A Romanized Bangla Hate Speech Corpus and a Systematic Benchmark of Parameter-Efficient LLM Adaptation
收藏资源简介:
RoBangHate is a manually annotated corpus for binary hate speech detection in Romanized Bangla (Banglish). The corpus consists of naturally occurring social-media comments collected from Facebook, Instagram, and YouTube. Each comment is annotated with one of two labels: Hate or Non-Hate. The annotation protocol distinguishes hate-directed content from ordinary criticism, political disagreement, negative sentiment, sarcasm, and profanity that does not constitute hate speech. The corpus is designed to capture linguistic characteristics of Banglish, including orthographic variation, informal spelling, Bangla-English code-mixing, profanity, and lexical diversity. The dataset is released to support research in Banglish hate speech detection, multilingual and low-resource NLP, content moderation, and robustness evaluation. This release includes the dataset, annotation guidelines, and documentation describing the corpus and annotation procedure. The dataset should be used for research and responsible NLP development. Users should not use the resource to identify, target,harass, or otherwise harm individuals.



