HarmVideoBench
收藏资源简介:
HarmVideoBench是一个由多所高校与企业联合构建的多层诊断性基准数据集,旨在评估大型多模态模型对有害视频的深度理解能力。该数据集包含1,379个视频及对应的4,137道多项选择题,数据来源于HateMM等公开资源以及从YouTube和Bilibili平台人工筛选的视频,涵盖多种有害场景与跨语言语境。数据集的构建采用了AI辅助与人工审核相结合的标注流程,通过三个层次(可观察证据、片段内部含义、跨片段推理)系统性地解构有害视频的理解任务。该数据集主要应用于内容安全领域,旨在解决现有基准在评估模型对隐含危害和上下文推理能力方面的不足,推动更透明、全面的自动化内容审核研究。
HarmVideoBench is a multi-layer diagnostic benchmark dataset jointly constructed by multiple universities and enterprises, aiming to evaluate the deep comprehension capabilities of large multimodal models towards harmful videos. This dataset includes 1,379 videos and their corresponding 4,137 multiple-choice questions, with data sourced from public resources such as HateMM and videos manually screened from YouTube and Bilibili platforms, covering various harmful scenarios and cross-lingual contexts. The dataset construction adopts an annotation workflow combining AI assistance and manual review, and systematically deconstructs the comprehension task of harmful videos through three levels: observable evidence, intra-segment meaning, and cross-segment reasoning. This dataset is mainly applied in the field of content security, aiming to address the shortcomings of existing benchmarks in evaluating models' abilities of implicit harm detection and contextual reasoning, and promote more transparent and comprehensive research on automated content moderation.
数据集概述:HarmVideoBench
来源与链接
- 论文地址:arXiv:2606.27187v1
- 提交日期:2026年6月25日
核心目标
- 评估大型多模态模型(LVLMs)对有害视频的深度理解能力,解决现有基准存在的两大问题:
- 忽视有害视频的多层次特征,仅将其简化为二分类任务,无法捕捉隐含或深层上下文危害。
- 缺乏解释性理由,仅判断模型是否正确标记,无法解释原因,导致评估黑箱化。
数据集构成
- 视频数量:1379个视频
- 题目数量:4137个多项选择题
- 评估维度:三个层次化维度
- 可观察证据 (Observable Evidence):基于视频表面信息进行判断。
- 片段-内部含义 (Clip-Internal Meaning):理解视频片段内部的深层语义。
- 超越片段推理 (Beyond-Clip Reasoning):需结合片段外部上下文进行推理。
评估与基线方法
- 评估模型:19个主流模型
- 提出的方法:BCR(基准对齐方法),可预测推理边界并仅在需要时动态检索上下文。
- 方法效果:将基础模型的宏平均准确率从61.7%提升至84.4%(达到当前最优)。
学科领域
- 计算机视觉与模式识别 (cs.CV)
- 计算与语言 (cs.CL)

- 1HarmVideoBench: Benchmarking Harmful Video Understanding in Large Multimodal Models中南大学; 清华大学; 华南师范大学; 字节跳动; 萨拉戈萨大学; CosmosMind; 武汉大学; 加州大学洛杉矶分校; 东南大学; 腾讯; 南开大学; Supermicro Computer Inc; 华中科技大学 · 2026年



