VideoMarathon
收藏资源简介:
VideoMarathon是一个大规模的时长视频指令跟随数据集,包含约9700小时的长时间视频,视频时长从3分钟到1小时不等。该数据集包含3.3M个高质量的QA对,涵盖了六个基本主题:时间性、空间性、对象、动作、场景和事件。与现有的视频指令数据集相比,VideoMarathon显著地扩展了训练视频的时长,支持22个多样化的任务,需要短期的和长期的视频理解。数据集的创建过程包括使用Qwen2VL-7B和DeepSeek-V3进行分层视频字幕生成,然后基于这些字幕合成QA对。VideoMarathon旨在解决现有视频语言模型在处理长时间视频时的长期依赖学习问题,支持更广泛的视频理解任务。
VideoMarathon is a large-scale long-form video instruction-following dataset, which contains approximately 9700 hours of long-duration videos ranging from 3 minutes to 1 hour in length. It includes 3.3 million high-quality QA pairs covering six core topics: temporality, spatiality, objects, actions, scenes, and events. Compared with existing video instruction datasets, VideoMarathon significantly extends the duration of training videos and supports 22 diverse tasks that require both short-term and long-term video understanding. The dataset creation process uses Qwen2VL-7B and DeepSeek-V3 to generate hierarchical video captions, then synthesizes QA pairs based on these captions. VideoMarathon aims to solve the long-term dependency learning problem of current video-language models when processing long-form videos, and supports a broader range of video understanding tasks.
VideoMarathon 数据集概述
数据集基本信息
- 名称: VideoMarathon
- 类型: 长视频指令跟随数据集
- 总时长: 约9,700小时
- 视频数量: 单个视频时长3至60分钟
- QA对数量: 3.3M高质量问答对
- 数据来源: 多样化视频领域
数据集特点
- 覆盖主题: 6个基础主题
- 时间性
- 空间性
- 对象
- 动作
- 场景
- 事件
- 任务类型: 22种多样化任务
- 支持短期和长期视频理解
- 比较优势:
- 显著延长训练视频时长(最长1小时)
- 更广的持续时间范围(3-60分钟)
- 更大规模的QA对数量
数据集构成
- 标注方式:
- 问题生成: Qwen2VL-7B
- 摘要生成: DeepSeek-V3
- 问题类型:
- 开放式问题(OE): 1.73M
- 多项选择题(MC): 1.57M
相关模型
- 模型名称: Hour-LLaVA
- 模型特点:
- 支持小时级视频训练和推理
- 1-FPS采样率
- 包含三个关键模块:
- 视频编码器
- 记忆增强模块(MemAug)
- LLM解码器
- 性能表现:
- 在3B和7-8B模型规模类别中
- 在TempCompass、LongVideoBench、Video-MME和LVBench四个基准测试上均取得最佳性能
引用信息
bibtex @article{lin2025unleashing, author = {Lin, Jingyang and Wu, Jialian and Sun, Ximeng and Wang, Ze and Liu, Jiang and Chen, Hao and Luo, Jiebo and Liu, Zicheng and Barsoum, Emad}, title = {Unleashing Hour-Scale Video Training for Long Video-Language Understanding}, journal = {arXiv preprint arXiv:2506.05332}, year = {2025}, }




