遇见数据集

BanglaCAT: A Context-Aware Bangla Social Media Toxicity Dataset with Hierarchical Annotation

收藏
Mendeley Data2026-09-08 收录
官方服务:

资源简介:

BanglaCAT is a context-aware toxicity dataset for Bangla (Bengali) social-media text. Each instance is a conversational pair consisting of a preceding comment (context) and a target comment whose toxicity is judged in light of that context. Unlike most existing Bangla resources that label isolated comments, this dataset is designed so that conversational context can both amplify and mitigate perceived toxicity (e.g., sarcasm, veiled threats, toxic endorsement, consensual banter, quoted or reported speech). Partitions: Train (19,453 pairs): LLM-generated candidates, cross-validated by independent LLMs, and human spot-checked. Released in original Bangla and English. Test (2,329 pairs): Real comments collected from public Facebook, YouTube, and Bangladeshi news-portal threads. Triple-annotated by human annotators; gold labels obtained by exact 2-of-3 joint hierarchical majority vote. Released in Bangla and English (row-aligned). Label schema (hierarchical): Level-1: Toxic / Non-toxic Level-2 (toxicity_type): None (when Non-toxic), or one of 11 fine-grained types when Toxic — Insult, Threat, Hate Speech, Profanity, Harassment, Identity Attack, Toxic Endorsement / Agreement, Incitement / Encouragement of Harm, Mockery / Ridicule, Sexual Harassment / Objectification, Dehumanization. The gold label is always the complete pair (level_1_class, toxicity_type). Level-1 and Level-2 are never voted independently, preventing invalid combinations. ## Dataset structure ``` BanglaCAT_Mendeley/ ├── train/ │ ├── bangla_context_toxicity_train.csv # Bangla (19,453 rows) │ └── bangla_context_toxicity_train_ENGLISH.csv # English translation, row-aligned ├── test/ │ ├── bangla_toxicity_test_set_HIERARCHICAL_FINAL.csv # Bangla (2,329 rows) │ └── bangla_toxicity_test_set_HIERARCHICAL_FINAL_ENGLISH.csv # Bangla + English columns ├── docs/ │ └── DATA_DICTIONARY.md ├── CITATION.cff ├── LICENSE.md ├── CHECKSUMS.txt └── README.md Train and test are strictly disjoint. Intended use is research on content moderation, hate-speech detection, and conversational AI safety.

创建时间:
2026-09-03
二维码
社区交流群
二维码
科研交流群
商业服务