VideoTemp-Bench
收藏资源简介:
VideoTemp-o3 数据集旨在支持代理式视频理解与推理任务,其核心是协调时间定位与视频理解。该数据集用于训练和评估能够处理视频问答的模型。给定一个视频及其对应的问题,模型需要执行按需的时间定位,以在长视频中找出与问题最相关的片段,并通过迭代优化该定位,最终基于定位到的关键视觉证据生成可靠的答案。数据集构建可能整合或借鉴了多个现有视频理解数据集,如 NExT-GQA、LongVideo-Reason、LongVILA 和 ScaleLong,以促进在复杂、长视频场景下的时序感知问答能力。
The VideoTemp-o3 dataset aims to support agent-based video understanding and reasoning tasks, with its core focus on coordinating temporal localization and video understanding. It is designed for training and evaluating models capable of handling video question answering. Given a video and its corresponding question, the model needs to perform on-demand temporal localization to identify the most relevant segments in long videos, iteratively optimize this localization, and ultimately generate reliable answers based on the localized key visual evidence. The dataset construction may integrate or draw from multiple existing video understanding datasets, such as NExT-GQA, LongVideo-Reason, LongVILA, and ScaleLong, to enhance temporal-aware question-answering capabilities in complex, long-video scenarios.





