遇见数据集

Moroccan Darija Offensive Language Detection Dataset

收藏
Mendeley Data2026-04-18 收录
官方服务:

资源简介:

The Moroccan Darija dataset was cleaned by removing duplicate entries and discarding sentences with conflicting annotations. To address class imbalance, undersampling was applied to reduce the size of the majority (non-offensive) class. The dataset was also augmented with samples from the OMCD corpus, which underwent the same preprocessing pipeline to ensure consistency, including emoji representation, normalization, removal of punctuation and diacritics, elimination of social media elements, elongation removal, and duplicate removal. Finally, all entries from the Moroccan Darija dataset were relabeled using Claude 3.5 Sonnet to align with the comprehensive OMCD framework, covering both explicit and implicit forms of offensiveness such as vulgarity, hate speech, hostile intent, contempt, humiliation, and belittlement (OMCD reference). Sentences where Claude-generated labels conflicting with previous annotations were flagged for manual review according to break the tie.

摩洛哥达里贾语(Moroccan Darija)数据集已通过移除重复条目、剔除标注冲突语句的方式完成清洗。为解决类别不平衡问题,我们对占比占优的非冒犯性类别实施欠采样以缩减其规模。 本数据集还通过引入OMCD语料库的样本完成扩充,且该部分样本经过了与主数据集一致的预处理流程以保证一致性,预处理步骤涵盖表情符号保留、文本归一化、标点与变音符号移除、社交媒体元素剔除、字符拉长现象去除以及重复内容清理。 最后,本数据集的所有条目均通过Claude 3.5 Sonnet进行重新标注,以适配涵盖显性与隐性冒犯形式的完整OMCD框架,冒犯形式包括低俗表述、仇恨言论、敌意意图、轻蔑、羞辱与贬低(OMCD参考标准)。 针对Claude生成的标注与原有标注存在冲突的语句,我们将其标记为需人工复核,以解决标注分歧。

创建时间:
2025-09-23
二维码
社区交流群
二维码
科研交流群
商业服务