遇见数据集

RoBangHate: A Romanized Bangla Hate Speech Corpus and a Systematic Benchmark of Parameter-Efficient LLM Adaptation

收藏
Zenodo2026-08-16 更新2026-08-20 收录
官方服务:

资源简介:

RoBangHate is a manually annotated corpus for binary hate speech detection in Romanized Bangla (Banglish). The corpus consists of naturally occurring social-media comments collected from Facebook, Instagram, and YouTube. Each comment is annotated with one of two labels: Hate or Non-Hate. The annotation protocol distinguishes hate-directed content from ordinary criticism, political disagreement, negative sentiment, sarcasm, and profanity that does not constitute hate speech. The corpus is designed to capture linguistic characteristics of Banglish, including orthographic variation, informal spelling, Bangla-English code-mixing, profanity, and lexical diversity. The dataset is released to support research in Banglish hate speech detection, multilingual and low-resource NLP, content moderation, and robustness evaluation. This release includes the dataset, annotation guidelines, and documentation describing the corpus and annotation procedure. The dataset should be used for research and responsible NLP development. Users should not use the resource to identify, target,harass, or otherwise harm individuals.

提供机构:
Zenodo
创建时间:
2026-08-16
二维码
社区交流群
二维码
科研交流群
商业服务