HourVideo
收藏资源简介:
HourVideo是一个用于长时间视频语言理解的基准数据集。它包含了一个新颖的任务套件,包括总结、感知(回忆、跟踪)、视觉推理(空间、时间、预测、因果、反事实)和导航(房间到房间、物体检索)任务。HourVideo包括从Ego4D数据集中手动挑选的500个以自我为中心的视频,持续时间为20到120分钟,并具有12,976个高质量的五路多项选择题。基准测试结果显示,多模态模型(包括GPT-4和LLaVA-NeXT)在随机机会上取得了微小的改进。相比之下,人类专家显著优于最先进的长时间上下文多模态模型Gemini Pro 1.5(85.0%对37.3%),突显了多模态能力上的巨大差距。我们希望将HourVideo建立为一个基准挑战,以推动能够真正理解无尽视觉数据流的先进多模态模型的发展。
HourVideo is a benchmark dataset for long-form video language understanding. It encompasses a novel task suite, including summarization, perception (recall, tracking), visual reasoning (spatial, temporal, prediction, causal, counterfactual), and navigation (room-to-room, object retrieval) tasks. HourVideo comprises 500 egocentric videos manually selected from the Ego4D dataset, with durations ranging from 20 to 120 minutes, and 12,976 high-quality five-way multiple-choice questions. Benchmark results show that multimodal models (including GPT-4 and LLaVA-NeXT) achieve only marginal improvements over random chance. By contrast, human experts significantly outperform the state-of-the-art long-form context multimodal model Gemini Pro 1.5 (85.0% vs 37.3%), highlighting the substantial gap in multimodal capabilities. We aim to establish HourVideo as a benchmark challenge to advance the development of advanced multimodal models that can truly comprehend endless visual data streams.
HourVideo: 1-Hour Video-Language Understanding
概述
HourVideo 是一个用于长时间视频语言理解的数据集,包含 500 个从 Ego4D 数据集中手动筛选的以自我为中心的视频,时长从 20 分钟到 120 分钟不等。数据集包含 12,976 个高质量的五选一多选题,涵盖总结、感知(回忆、跟踪)、视觉推理(空间、时间、预测、因果、反事实)和导航(房间到房间、物体检索)任务。
数据集组成
- 视频数量: 500 个
- 视频时长: 20 分钟到 120 分钟
- 问题数量: 12,976 个五选一多选题
任务类型
- 总结
- 感知
- 回忆
- 跟踪
- 视觉推理
- 空间
- 时间
- 预测
- 因果
- 反事实
- 导航
- 房间到房间
- 物体检索
基准结果
- GPT-4: 平均得分 19.6%
- LLaVA-34B-DPO: 平均得分 22.3%
- Gemini 1.5 Pro: 平均得分 37.3%
数据集下载
- 开发集: 包含 50 个视频,1182 个多选题,时长 39.3 小时。下载地址:HourVideo 开发集
联系信息
- Keshigeyan Chandrasegaran: keshik@stanford.edu
- Agrim Gupta: agrim@stanford.edu
- Lea M. Hadzic: lea27@stanford.edu
- Manling Li: manlingl@stanford.edu
引用
bibtex @inproceedings{chandrasegaran2024hourvideo, title={HourVideo: 1-Hour Video-Language Understanding}, author={Chandrasegaran, Keshigeyan and Gupta, Agrim and Hadzic, Lea M. and Kota, Taran and He, Jimming and Eyzaguirre, Cristobal and Durante, Zane and Li, Manling and Wu, Jiajun and Li, Fei-Fei}, booktitle = {Advances in Neural Information Processing Systems}, year={2024}, volume = {37}, }




