VideoHallu
收藏资源简介:
VideoHallu是一个基准测试集,由流行模型如Sora、Veo2、Kling生成的合成视频构建而成,并配有专家制作的问答对示例,这些示例可以轻松通过人类水平的感知和推理解决。该数据集旨在评估多模态大语言模型(MLLMs)在合成视频中检测异常内容的能力。
VideoHallu is a benchmark dataset constructed from synthetic videos generated by popular models such as Sora, Veo2, and Kling, accompanied by question-answer pairs crafted by experts, which can be readily resolved through human-level perception and reasoning. This dataset is designed to evaluate the capability of multi-modal large language models (MLLMs) in detecting anomalous content within synthetic videos.
VideoHallu 数据集概述
数据集基本信息
- 名称: VideoHallu
- 发布日期: 2025年5月2日
- 数据集大小: 3233个样本
- 存储位置: HuggingFace
- 相关论文: VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations for Synthetic Videos
数据集简介
VideoHallu是一个用于评估和缓解合成视频中多模态幻觉的基准数据集。该数据集由流行的视频生成模型(如Sora、Veo2、Kling)生成的合成视频组成,并配有专家精心设计的问答对示例。
数据集类别
数据集包含四个主要问题类别:
- Alignment: 检查模型是否正确识别和理解实体。
- Spatial-temporal Consistency: 检查模型是否能跟踪实体在帧间的运动。
- Common Sense Reasoning: 测试模型基于知识的推理能力。
- Physics: 评估模型对物理定律的应用能力。
数据集结构
- 数据格式: JSON文件
- 主要文件:
- synthetic_data_split.json
- physbench_train_split.json
评估模型
数据集评估了多个先进的MLLMs,包括:
- GPT-4o
- Gemini-2.5-Pro
- Qwen-2.5-VL
- Video-R1
- VideoChat-R1
训练与微调
- 训练方法: 使用Group Relative Policy Optimization (GRPO)对Qwen-2.5-VL-7B进行微调。
- 训练数据: 包含真实世界和合成的常识/物理数据集。
- 结果: 微调后的模型在合成视频理解方面表现更优。
奖励模型
- 基础模型: ModernBERT
- 微调数据集: MOCHA, Prometheus-preference, Pedants
- 功能: 评估自由形式的文本生成。
引用
如需使用该数据集,请引用相关论文: bibtex @misc{li2025videohalluevaluatingmitigatingmultimodal, title={VideoHallu: Evaluating and Mitigating Multi-modal Hallucinations for Synthetic Videos}, author={Zongxia Li and Xiyang Wu and Yubin Qin and Guangyao Shi and Hongyang Du and Dinesh Manocha and Tianyi Zhou and Jordan Lee Boyd-Graber}, year={2025}, eprint={2505.01481}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2505.01481}, }




