Video MindPalace Benchmark (VMB)
收藏资源简介:
Video MindPalace Benchmark (VMB) 是一个用于评估模型在真实环境中进行空间、时间和布局关系推理能力的新型基准测试。该数据集由威斯康星大学麦迪逊分校、Meta和伊利诺伊大学厄巴纳-香槟分校的研究团队创建,旨在通过第一人称视角视频捕捉人类日常活动,并生成与3D世界紧密相关的数据。VMB包含三类问题:增强空间定位、上下文时间推理和布局感知推理,要求模型提供类似于人类理解的上下文响应。该数据集的应用领域主要集中在长视频理解和大规模视觉语言模型的推理能力提升上,旨在解决长视频分析中的时空一致性和人类对齐推理问题。
Video MindPalace Benchmark (VMB) is a novel benchmark for evaluating models' capacity to reason about spatial, temporal, and layout relationships in real-world scenarios. This dataset was developed by research teams from the University of Wisconsin-Madison, Meta, and the University of Illinois Urbana-Champaign, which aims to capture human daily activities via first-person perspective videos and generate data closely associated with the 3D world. VMB includes three types of questions: enhanced spatial localization, contextual temporal reasoning, and layout-aware reasoning, requiring models to deliver contextual responses analogous to human comprehension. The application fields of this dataset mainly focus on long-form video understanding and enhancing the reasoning capabilities of large-scale vision-language models, with the goal of addressing the issues of spatio-temporal consistency and human-aligned reasoning in long-form video analysis.




