etri-vilab/MultihopSpatial
收藏资源简介:
MultihopSpatial 是一个基准测试数据集,旨在评估视觉语言模型(VLMs)在多跳组合空间推理方面的鲁棒性。与仅评估单步空间关系的现有基准不同,MultihopSpatial 包含1到3个推理跳数的查询,并配有视觉定位评估,揭示了一个关键盲点:在多项选择题中准确率高的模型往往缺乏正确的空间定位能力。所有4,500个基准QA对和边界框均由十名训练有素的人类专家严格标注,内部评分者一致性为90%(Krippendorffs α = 0.90)。该数据集支持自我中心和外中心视角,涵盖属性、位置和关系三种空间类别,并可组合成多跳问题。此外,还提供训练数据MultihopSpatial-Train(6,791个样本),支持通过强化学习进行后训练。
MultihopSpatial is a benchmark designed to evaluate whether vision-language models (VLMs) demonstrate robustness in multi-hop compositional spatial reasoning. Unlike existing benchmarks that only assess single-step spatial relations, MultihopSpatial features queries with 1 to 3 reasoning hops paired with visual grounding evaluation, exposing a critical blind spot: models achieving high multiple-choice accuracy often lack proper spatial localization. All 4,500 benchmark QA pairs and bounding boxes are strictly annotated by ten trained human experts with an inter-rater agreement of 90% (Krippendorffs α = 0.90). It includes perspective-taking (ego-centric and exo-centric viewpoints) and three spatial categories (Attribute, Position, and Relation) composable into multi-hop questions. The dataset also provides training data, MultihopSpatial-Train (6,791 samples), to support post-training via reinforcement learning.




