ShortVideo Multimodal Adversarial (SVMA) 数据集
收藏资源简介:
SVMA数据集是一个包含1009个短视频的全面基准数据集,旨在评估现代多模态大型语言模型的内容安全性。该数据集包含了丰富多样的短视频内容及其人类引导的合成对抗性攻击,是首个专为短视频内容审核设计的多模态对抗性数据集。数据集涵盖了多种文化、主题和感知边界的内容,包括讽刺、社会评论、个人故事和合成虚假信息等。数据集采用了严格的原则性定义来确定内容是否适当,并根据适当的程度将内容分为不适当和适当两类。SVMA数据集是评估多模态推理失败的高保真沙盒。
The SVMA dataset is a comprehensive benchmark dataset containing 1,009 short videos, designed to evaluate the content safety of modern multimodal large language models. This dataset features a rich variety of short video content and human-guided synthetic adversarial attacks, and it is the first multimodal adversarial dataset specifically designed for short video content moderation. The dataset covers content spanning diverse cultures, topics, and perceptual boundaries, including satire, social commentary, personal narratives, and synthetic disinformation, among others. The dataset adopts strict principled definitions to determine whether content is appropriate, and categorizes content into inappropriate and appropriate categories based on the degree of appropriateness. The SVMA dataset is a high-fidelity sandbox for evaluating multimodal reasoning failures.
数据集概述
基本信息
- 数据集名称: SVMA (Short-Video Multimodal Adversarial) dataset
- 相关论文: Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation
- 会议: ICCV 2025 SVU workshop
- 论文链接: arXiv:2507.11968
- 代码仓库: ChimeraBreak
数据集内容
- 用途: 用于评估短格式视频内容审核的对抗性数据集
- 数据类型: 多模态(视频、音频、文本)
- 攻击类型: 三模态对抗攻击(视觉、听觉、文本)
获取方式
- HuggingFace: SVMA-dataset
- Kaggle: SVMA-bench
相关工具
- 代码库结构:
data/: 包含标注和HF管道脚本notebooks/: 包含攻击和评估笔记本及评估指标utils/: 包含标注提示和合成标注脚本
引用信息
bibtex @misc{mustakim2025watchlistenunderstandmislead, title={Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation}, author={Sahid Hossain Mustakim and S M Jishanul Islam and Ummay Maria Muna and Montasir Chowdhury and Mohammed Jawwadul Islam and Sadia Ahmmed and Tashfia Sikder and Syed Tasdid Azam Dhrubo and Swakkhar Shatabda}, year={2025}, eprint={2507.11968}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2507.11968}, }
许可证
- 许可证类型: 未明确说明(需查看LICENSE文件)

- 1Watch, Listen, Understand, Mislead: Tri-modal Adversarial Attacks on Short Videos for Content Appropriateness Evaluation联合国际大学, BRAC大学, 不列颠哥伦比亚大学, 孟加拉国专业大学, 阿尔伯塔大学 · 2025年



