Harm or Humor Benchmark
收藏资源简介:
该数据集由穆罕默德·本·扎耶德人工智能大学团队构建,是一个多模态、多语言的幽默有害性检测基准,涵盖英语和阿拉伯语的3000条文本、6000张图像及1200段视频。数据通过人工严格标注,区分安全笑话与有害笑话(显性/隐性),重点捕捉文化语境中的隐性伤害。其创新性在于融合低资源语言与跨模态内容,用于评估AI模型对复杂文化推理任务的处理能力,推动安全对齐技术发展。
This dataset is a multimodal, multilingual benchmark for humorous harm detection, developed by a team from Mohamed bin Zayed University of Artificial Intelligence (MBZUAI). It comprises 3,000 text samples, 6,000 images, and 1,200 video clips in both English and Arabic. All data undergoes strict manual annotation to categorize content into safe jokes and harmful jokes (including both explicit and implicit forms), with a particular emphasis on capturing subtle harms rooted in cultural contexts. The novelty of this benchmark lies in its integration of low-resource languages and cross-modal content, which enables the evaluation of AI models' performance on complex cultural reasoning tasks and facilitates the advancement of safety alignment technologies.

- 1Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor穆罕默德·本·扎耶德人工智能大学; 索非亚大学·INSAIT研究所 · 2026年



