遇见数据集

QCRI/ArGuard-Task1

收藏
Hugging Face2026-05-24 更新2026-06-14 收录
官方服务:

资源简介:

ArGuard – Track A (Arabic Hateful Memes) 数据集是ArGuard共享任务中Track A的官方数据集,专注于阿拉伯语仇恨迷因的多模态检测。每个数据实例包括一个阿拉伯语迷因(图像和通过OCR提取的覆盖文本),并手动标注了仇恨性(二元分类:仇恨或非仇恨)和细粒度子类型(多标签分类,如嘲笑、煽动、非人化、侮辱性用语、蔑视、低劣性、排斥等仇恨子类型,以及幽默、讽刺等非仇恨子类型)。数据集包含训练集(3,500条记录)、开发集(500条记录)、开发测试集(500条记录,标签被移除)和测试集(500条记录,三重标注)。数据集旨在用于阿拉伯语多模态仇恨言论检测的研究,包括二元分类和细粒度预测,但包含冒犯性内容,仅限非商业研究使用。

ArGuard – Track A (Arabic Hateful Memes) dataset is the official dataset for Track A of the ArGuard shared task, focusing on multimodal hateful-meme detection in Arabic. Each instance consists of an Arabic meme (image and OCR-extracted overlaid text) manually annotated for hatefulness (binary classification: Hateful or Not Hateful) and fine-grained sub-types (multi-label classification, such as Mocking, Incitement, Dehumanization, Slurs, Contempt, Inferiority, Exclusion for hateful sub-types, and Humor, Sarcasm for non-hateful sub-types). The dataset includes train (3,500 records), dev (500 records), dev_test (500 records with labels dropped), and test (500 records, triple-annotated) splits. It is intended for research on Arabic multimodal hate speech detection, including binary classification and fine-grained prediction, but contains offensive content and is for non-commercial research use only.

提供机构:
QCRI
二维码
社区交流群
二维码
科研交流群
商业服务