HLV-1K
收藏资源简介:
HLV-1K是由抖音、南洋理工大学等机构联合创建的大规模长时间视频理解基准数据集,旨在评估模型在长时间视频内容上的理解能力。该数据集包含1009个时长超过一小时的视频,总计14,847个高质量的问题回答对,涵盖了帧级、事件内级、跨事件级和长期推理任务。数据集的创建过程包括视频收集、关键帧提取、事件标注以及问题生成等多个步骤,确保了数据的多样性和高质量。HLV-1K的应用领域主要集中在长时间视频理解任务,如直播视频、会议记录和电影等,旨在解决长时间视频内容中的复杂时空关系理解和长期依赖性问题。
HLV-1K is a large-scale long-duration video understanding benchmark dataset jointly developed by institutions including Douyin and Nanyang Technological University, et al. It is designed to evaluate the video understanding capabilities of models on long-form video content. This dataset comprises 1009 videos each with a duration of over one hour, totaling 14,847 high-quality question-answer pairs, covering frame-level, intra-event, cross-event, and long-term reasoning tasks. The construction of HLV-1K involves multiple sequential steps: video collection, key frame extraction, event annotation, and question generation, which ensures the dataset's diversity and high data quality. The primary application scenarios of HLV-1K are long-duration video understanding tasks, including live streaming videos, meeting recordings, and movies, aiming to tackle the challenges of complex spatio-temporal relationship understanding and long-term dependency in long-form video content.

- 1HLV-1K: A Large-scale Hour-Long Video Benchmark for Time-Specific Long Video Understanding抖音、南洋理工大学、西南交通大学、大湾区大学、深圳大学 · 2025年



