Code-mixed Chaos : Multi-labeled Banglish & Bangla Corpus for Toxicity analysis
收藏资源简介:
The dataset addresses a crucial gap in toxicity detection for Banglish—a code-mixed form of Bengali and English written in Roman script—which is often undervalued in NLP research. To mitigate this, we present a manually collected, multi-labeled dataset comprising 10,234 Banglish social media comments, annotated across 10 classes with toxic and non-toxic categories. The toxic comments are categorized into nine types: (1) Vulgar-based, (2) Religious-Hostility, (3) Troll-based, (4) Insult-based, (5) Loathe-based, (6) Threat-based, (7) Race-based, (8) Sexual-based, and (9) Political-Chaos. And a single Non-toxic category representing comments that do not have any form of toxicity. It is equally divided between toxic (5,117) and non-toxic (5,117) entries. Each sample was sourced from platforms such as Facebook, YouTube, Instagram, and X (formerly Twitter). To balance the dataset, it is enriched by selectively adding non-toxic texts from a publicly available corpus: "Bengali & Banglish: A Monolingual Dataset for Emotion Detection in Linguistically Diverse Contexts". Additionally, we provided a Bangla-translated version of the dataset to support the script-based comparative analysis in toxicity detection.
本数据集针对孟加拉英语混合语(Banglish,一种以罗马字母书写的孟加拉语与英语代码混合语)的毒性检测任务长期存在的关键研究空白——该语言在自然语言处理(Natural Language Processing,NLP)研究中常被低估。为填补这一空白,我们构建了一套人工采集的多标签数据集,包含10234条孟加拉英语混合语社交媒体评论,共标注为10个类别,涵盖有毒与无毒两大分类。其中有毒评论被划分为9个子类:(1) 低俗类(Vulgar-based)、(2) 宗教敌对类(Religious-Hostility)、(3) 引战类(Troll-based)、(4) 侮辱类(Insult-based)、(5) 恶意厌恶类(Loathe-based)、(6) 威胁类(Threat-based)、(7) 种族类(Race-based)、(8) 性相关类(Sexual-based)以及(9) 政治混乱类(Political-Chaos);另有一个独立的无毒类别,用于指代未携带任何形式毒性的评论。数据集的有毒评论与无毒评论数量均等,各为5117条。所有样本均采集自Facebook、YouTube、Instagram及X(原Twitter)平台。为平衡数据集分布,我们从公开可用的语料库《Bengali & Banglish: A Monolingual Dataset for Emotion Detection in Linguistically Diverse Contexts》中选择性添加无毒文本以扩充数据集。此外,我们还提供了该数据集的孟加拉语翻译版本,以支持基于书写脚本的毒性检测对比分析。



