遇见数据集

ArabXlur: An Arabic Taboo-Language Corpus for Detecting Obfuscated Sexual-Taboo Language (v1.0)

收藏
Zenodo2026-07-16 更新2026-08-01 收录
官方服务:

资源简介:

ArabXlur is a manually annotated corpus of 30,150 Arabic tweets built to study the detection of sexually taboo language that has been orthographically camouflaged to evade automated moderation. In Arabic, dots and diacritics carry phonemic weight and therefore double as instruments of visual disguise (for example, swapping a dotted letter for its dotless placeholder, or padding a word with diacritics). The corpus treats such disguise as a labeled class in its own right. The dataset is balanced across three classes: clean text (label 0), overt taboo in canonical spelling (label 1), and orthographically camouflaged taboo (label 2), 10,050 tweets each. It is released with an 80/20 stratified train/test split (train 24,120; test 6,030) and a 126-item seed lexicon in canonical and disguised forms. Full documentation is in the included README. ACCESS IS RESTRICTED. This record contains explicit and offensive language and a lexicon of sexually taboo terms, released solely as a moderation-research aid. Because of the residual dual-use risk, access is granted only to named researchers for non-commercial academic use, under the conditions stated for this record. To request access, use the "Request access" button and confirm that you accept those conditions; requests are reviewed and approved individually by the provider. Associated article: Alqarni, M. (forthcoming). Camouflage as a Class: An Arabic Taboo-Language Corpus and a Test of Learnability Beyond the Lexicon.

提供机构:
Zenodo
创建时间:
2026-07-16
二维码
社区交流群
二维码
科研交流群
商业服务