Yongxin-Guo/VTG-IT
收藏资源简介:
VTG-IT-120K是一个高质量且全面的指令调优数据集,涵盖了视频时间定位(VTG)任务,包括时刻检索(63.2K)、密集视频字幕(37.2K)、视频摘要(15.2K)和视频亮点检测(3.9K)。此外,VTG-LLM模型有效地将时间戳知识整合到视觉令牌中,并引入了绝对时间令牌来处理时间戳知识,避免了概念转移,并引入了一种轻量级、高性能的基于槽的令牌压缩方法,以促进更多视频帧的采样。
VTG-IT-120K is a high-quality and comprehensive instruction tuning dataset that covers VTG tasks such as moment retrieval (63.2K), dense video captioning (37.2K), video summarization (15.2K), and video highlight detection (3.9K). Additionally, the VTG-LLM model effectively integrates timestamp knowledge into visual tokens, incorporates absolute-time tokens that specifically handle timestamp knowledge, thereby avoiding concept shifts, and introduces a lightweight, high-performance slot-based token compression method to facilitate the sampling of more video frames.
数据集概述
数据集名称
- VTG-IT-120K
数据集内容
- 包含多种视频时间定位任务的数据,具体包括:
- 时刻检索(63.2K)
- 密集视频字幕(37.2K)
- 视频摘要(15.2K)
- 视频亮点检测(3.9K)
数据集特点
- 高质量和全面性
- 用于指令调优
VTG-LLM技术特点
- 有效整合时间戳知识到视觉令牌中
- 引入绝对时间令牌,专门处理时间戳知识,避免概念偏移
- 采用轻量级、高性能的基于槽的令牌压缩方法,便于采样更多视频帧
引用信息
@article{guo2024vtg, title={VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding}, author={Guo, Yongxin and Liu, Jingyu and Li, Mingda and Tang, Xiaoying and Chen, Xi and Zhao, Bo}, journal={arXiv preprint arXiv:2405.13382}, year={2024} }




