Bangla Abusive Regional Dialect Dataset
收藏资源简介:
BanglaDial-Abuse is a balanced Bengali-script dataset developed for regional dialect identification in abusive and hostile Bangla text. The dataset contains 1,000 sentences distributed equally across four linguistic varieties: Standard Bangla, Chattagram, Sylhet, and Barishal, with 250 samples per class. The primary research task is multi-class regional dialect classification. Given an abusive or hostile Bangla sentence, the objective is to identify the regional variety represented by the text. The dataset is therefore intended for dialect identification rather than binary abusive-language detection, as all samples belong to the abusive or hostile-language domain. The dataset was developed as a corpus-grounded synthetic regional-language resource. Regional sentence construction was guided by recurring linguistic patterns observed in Bangla regional-language data, including differences in pronouns, possessive forms, verb morphology, negation, interrogative structures, postpositions, regional vocabulary, and Bengali-script spelling conventions. The construction process aimed to preserve the underlying hostile or abusive meaning while expressing it through linguistically characteristic regional forms. The dataset contains two main fields: text — abusive or hostile Bangla sentence dialect_label — corresponding regional variety The resource can support research in Bangla dialect identification, dialect-aware abusive-language NLP, low-resource NLP, cross-dialect robustness, regional lexical and morphological analysis, character-level versus token-level modeling, transfer learning, and explainable dialect classification. Potential baseline approaches include character and word-level TF-IDF models, Support Vector Machines, BanglaBERT, mBERT, and XLM-RoBERTa. Important Note This release should be treated as a research and prototyping corpus. Because the regional examples were synthetically constructed using corpus-derived linguistic patterns, some sentences may contain unnatural expressions, mixed regional features, orthographic inconsistencies, or highly distinctive dialect markers. Future versions are intended to incorporate systematic native-speaker validation and annotation. Ethical Considerations The dataset contains abusive and potentially offensive language and should be used primarily for legitimate research purposes, including language technology, dialect identification, moderation research, and model robustness evaluation. Regional dialect labels represent linguistic varieties in the dataset and should not be interpreted as evidence about the identity, behavior, or characteristics of individual speakers. Dataset size: 1,000 samplesNumber of classes: 4Samples per class: 250Language: Bangla/BengaliScript: BengaliTask: Multi-class regional dialect classificationVersion: 1.0



