TimeLens-100K
收藏资源简介:
TimeLens-100K是一个大规模、多样化且高质量的视频时间定位训练数据集。该数据集在我们的论文《TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs》中提出,并用于训练TimeLens模型。标注过程使用了由Gemini-2.5-Pro驱动的自动化流程。数据集统计信息包括:总视频数约20K,总标注数约100K,平均每个视频有5个标注。视频来源多样,包括DiDeMo、QuerYD、HiREST、CosMo-Cap和InternVid-VTime等数据集。
TimeLens-100K is a large-scale, diverse and high-quality video temporal grounding training dataset. It was proposed in our paper titled "TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs" and is used to train the TimeLens model. The annotation process adopts an automated pipeline powered by Gemini-2.5-Pro. Dataset statistics are as follows: there are approximately 20,000 total videos and around 100,000 total annotations, with an average of 5 annotations per video. The videos are sourced from various datasets including DiDeMo, QuerYD, HiREST, CosMo-Cap, InternVid-VTime and others.
TimeLens-100K 数据集概述
数据集基本信息
- 名称:TimeLens-100K
- 许可证:bsd-3-clause (https://github.com/TencentARC/TimeLens/blob/main/LICENSE)
- 主要语言:英语 (en)
- 任务类别:视频文本到文本 (video-text-to-text)
- 数据规模:10K<n<100K
数据集描述
TimeLens-100K 是一个用于视频时序定位的大规模、多样化、高质量训练数据集。该数据集在论文《TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs》中提出,并用于训练 TimeLens 模型。其标注过程采用由 Gemini-2.5-Pro 驱动的自动化流程完成。
数据集统计
- 视频总数:约 20K
- 标注总数:约 100K
- 平均每视频标注数:约 5
- 视频来源:视频采样自多个数据集:
- DiDeMo:https://github.com/LisaAnne/LocalizingMoments/
- QuerYD:https://www.robots.ox.ac.uk/~vgg/data/queryd/
- HiREST:https://github.com/j-min/HiREST
- CosMo-Cap:https://github.com/showlab/cosmo
- InternVid-VTime:https://github.com/OpenGVLab/InternVideo/tree/main/Data/InternVid
相关资源
- 论文:https://arxiv.org/abs/2512.14698
- 代码仓库:https://github.com/TencentARC/TimeLens
- 项目主页:https://timelens-arc-lab.github.io/
- 模型与数据集合:https://huggingface.co/collections/TencentARC/timelens
- 使用说明:下载和使用该数据集进行训练的指南请参考 GitHub 仓库 (https://github.com/TencentARC/TimeLens#-training-on-timelens-100k)。
引用
如需在研究中引用此数据集,请使用以下 BibTeX 条目: bibtex @article{zhang2025timelens, title={TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs}, author={Zhang, Jun and Wang, Teng and Ge, Yuying and Ge, Yixiao and Li, Xinhao and Shan, Ying and Wang, Limin}, journal={arXiv preprint arXiv:2512.14698}, year={2025} }




