VideoITG-40K
收藏资源简介:
VideoITG-40K是一个大规模的视频理解数据集,由香港理工大学、南京大学、英伟达和哈佛大学的研究人员构建。该数据集包含40,000个视频和500,000个指令引导的时间定位标注,旨在解决长视频理解中的复杂场景问题。VideoITG通过指令引导的帧采样,能够有效处理多时态线索,并针对不同任务需求定制帧选择策略。VidThinker是一个自动化数据标注流程,通过指令引导的剪辑字幕、检索和帧定位,确保了高质量的标注。VideoITG-40K数据集的创建过程借鉴了人类的推理过程,采用粗到精的策略,并使用GPT-4o进行详细剪辑描述,随后通过“Needle-In-A-Haystack”方法进行指令引导的剪辑检索。数据集的分类指令分为四类,分别对应视频问答中的不同推理需求。VideoITG-40K数据集在规模和质量上都显著超越了之前的时间定位数据集,为视频理解模型的训练提供了丰富的资源。
VideoITG-40K is a large-scale video understanding dataset constructed by researchers from The Hong Kong Polytechnic University, Nanjing University, NVIDIA, and Harvard University. This dataset contains 40,000 videos and 500,000 instruction-guided temporal localization annotations, aiming to address complex scene challenges in long-form video understanding. VideoITG adopts instruction-guided frame sampling, which can effectively handle multi-temporal cues and customize frame selection strategies for different task requirements. VidThinker is an automated data annotation pipeline that ensures high-quality annotations via instruction-guided clip captioning, retrieval, and frame localization. The creation process of the VideoITG-40K dataset draws on human reasoning procedures, adopting a coarse-to-fine strategy: it first uses GPT-4o to generate detailed clip descriptions, then conducts instruction-guided clip retrieval via the "Needle-In-A-Haystack" method. The dataset’s classification instructions are divided into four categories, corresponding to different reasoning demands in video question answering. The VideoITG-40K dataset significantly outperforms previous temporal localization datasets in both scale and quality, providing rich resources for the training of video understanding models.
VideoITG数据集概述
基本信息
- 数据集名称:VideoITG (Instructed Temporal Grounding for Videos)
- 开发团队:Shihao Wang等(香港理工大学、NVIDIA、南京大学、哈佛大学)
- 联系方式:cslzhang@comp.polyu.edu.hk; scutchrisding@gmail.com
- 相关资源:arXiv论文、代码模型(即将发布)、数据集
核心创新
- 提出指令引导的时序定位框架(VideoITG)
- 开发VidThinker自动化标注流程:
- 基于指令的片段级视频描述生成
- 指令引导的相关片段检索
- 细粒度帧级定位
数据集详情
- 名称:VideoITG-40K
- 规模:40,000个视频
- 标注量:500,000条指令时序定位标注
- 标注策略:
- 语义聚焦指令:选择包含关键视觉线索的多样化帧
- 运动聚焦指令:均匀采样以捕捉动态变化
- 混合需求:应用混合采样策略
- 开放指令:全视频最小多样化帧采样
模型设计
- 文本生成:对齐视频和语言token进行序列预测
- 分类架构:
- 因果注意力:使用锚点token管理时序线索
- 完全注意力:促进视觉和文本token的跨模态交互
性能表现
- 与不同Video-LLMs集成时获得持续性能提升
- 在多模态视频理解基准测试中表现优越
引用信息
bibtex @misc{wang2025videoitgmultimodalvideounderstanding, title={VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding}, author={Shihao Wang et al.}, year={2025}, eprint={2507.13353}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2507.13353} }
许可信息
- 网站授权:Creative Commons Attribution-ShareAlike 4.0 International License
- 1VideoITG: Multimodal Video Understanding with Instructed Temporal Grounding香港理工大学, 南京大学, 英伟达, 哈佛大学 · 2025年



