QCRI/MemeReason
收藏资源简介:
MemeReason是一个扩展的多模态数据集,旨在通过思维链监督进行可解释的仇恨和宣传性表情包检测。它基于两个基准数据集进行增强:hateful_memes(英语,二分类仇恨表情包检测)和armeme(阿拉伯语,4类宣传表情包检测)。数据集扩展了自然语言解释、细粒度标签(如受保护类别和攻击类型)、宣传技术注释(仅限armeme)以及从大型语言模型(如GPT-4.1)提取的逐步思维链推理。这些扩展用于训练可解释的、基于思维的多模态大语言模型,支持图像分类、视觉问答和文本生成等任务。数据集包含训练、开发和测试分割,适用于研究仇恨言论检测、宣传检测和可解释人工智能。注意:数据集内容可能包含令人不安或冒犯性的表情包。
MemeReason is an extended multimodal dataset for explainable detection of hateful and propagandistic memes with chain-of-thought supervision. It augments two benchmarks: hateful_memes (English, binary hateful meme detection) and armeme (Arabic, 4-class propaganda meme detection). The dataset includes natural-language explanations, fine-grained labels (e.g., protected category and attack type), propaganda technique annotations (for armeme), and step-by-step chain-of-thought rationales distilled from large language models like GPT-4.1. These extensions are used to train explainable, thinking-based multimodal LLMs, supporting tasks such as image classification, visual question answering, and text generation. The dataset contains train, dev, and test splits and is suitable for research on hate-speech detection, propaganda detection, and explainable AI. Warning: the dataset contains memes with potentially disturbing or offensive content.




