V-STaR
收藏资源简介:
V-STaR数据集是由伦敦玛丽女王大学、南京大学和南洋理工大学的研究人员构建的,旨在评估视频大型语言模型在视频时空推理方面的能力。该数据集包含由半自动化GPT-4驱动的管道生成的粗到细的CoT问题,模仿人类的认知过程,嵌入明确的推理链。数据集通过分解视频理解任务为逆时空推理任务,同时评估模型在识别对象、事件发生时间和对象位置方面的能力,以及模型构建CoT逻辑的过程。
V-STaR dataset was constructed by researchers from Queen Mary University of London, Nanjing University and Nanyang Technological University, with the aim of evaluating the capabilities of video large language models (LLMs) in video spatio-temporal reasoning. The dataset comprises coarse-to-fine Chain-of-Thought (CoT) questions generated via a semi-automated GPT-4-driven pipeline, which mimics human cognitive processes and embeds explicit reasoning chains. Furthermore, the dataset decomposes video understanding tasks into reverse spatio-temporal reasoning tasks, while simultaneously assessing the model's abilities in object recognition, event timing detection, object localization, as well as the process through which the model constructs CoT reasoning logic.

- 1V-STaR: Benchmarking Video-LLMs on Video Spatio-Temporal Reasoning伦敦玛丽女王大学, 南京大学, 南洋理工大学 · 2025年



