TimeLens-Bench, TimeLens-100K
收藏资源简介:
TimeLens-Bench是由南京大学、腾讯PCG等机构联合构建的高质量视频时序定位基准,包含对Charades-STA、ActivityNet Captions和QVHighlights三个流行数据集的重新标注版本,严格遵循事件唯一性、存在性和标注准确性等标准。该数据集规模达10万条,通过人工审核和自动化流程修正了原始数据中34.9%的标注错误,显著提升了数据的可靠性。数据集构建采用诊断-修正工作流,包括交叉验证和难度分级采样,主要应用于多模态大语言模型的时序感知能力训练,旨在解决视频理解中'何时发生'的精确定位问题。
TimeLens-Bench is a high-quality video temporal localization benchmark jointly constructed by institutions including Nanjing University and Tencent PCG. It includes re-annotated versions of three popular datasets: Charades-STA, ActivityNet Captions, and QVHighlights, which strictly adhere to standards such as event uniqueness, annotation existence, and annotation accuracy. The dataset consists of 100,000 annotated entries, with 34.9% of the annotation errors in the original data corrected via manual review and automated workflows, significantly improving the dataset's reliability. The construction of TimeLens-Bench adopts a diagnostic-correction workflow encompassing cross-validation and difficulty-stratified sampling. It is primarily used for training multimodal large language models on their temporal perception capabilities, aiming to address the precise localization problem of "when an event occurs" in video understanding.

- 1TimeLens: Rethinking Video Temporal Grounding with Multimodal LLMs南京大学, ARC Lab, 腾讯PCG, 上海AI Lab · 2025年



