Neptune
收藏资源简介:
Neptune是由谷歌研究团队创建的一个用于长视频理解的数据集,旨在解决现有数据集多集中于短视频片段的问题。该数据集包含3268个问题-答案-干扰项的标注,涵盖了2405个视频,视频长度从几秒到15分钟不等。数据集的创建过程利用了大规模的视频语言模型(VLMs)和大型语言模型(LLMs),自动生成时间对齐的视频字幕和复杂的问题-答案-干扰项集。Neptune特别强调多模态推理能力,适用于评估模型在长视频中的时间顺序、计数和状态变化等方面的表现,旨在推动更先进的长视频理解模型的开发。
Neptune is a long-form video understanding dataset developed by the Google Research team, which aims to address the limitation that existing datasets mostly focus on short video clips. This dataset contains 3268 question-answer-distractor annotations, covering 2405 videos with durations ranging from several seconds to 15 minutes. The development of this dataset leverages large-scale video language models (VLMs) and large language models (LLMs) to automatically generate time-aligned video captions and complex question-answer-distractor sets. Neptune places special emphasis on multimodal reasoning capabilities, and is suitable for evaluating model performance in aspects such as temporal order, counting, and state changes in long-form videos, aiming to promote the development of more advanced long-form video understanding models.




