Spatial-DISE
收藏资源简介:
Spatial-DISE是一个用于评估视觉语言模型(VLMs)空间推理能力的统一基准。该数据集由利物浦大学计算机科学系的研究团队创建,包含超过12K个经过验证的空间推理视觉问答(VQA)对。数据集分为Spatial-DISE Bench(559个评估VQA对)和Spatial-DISE-12K(12K+个训练VQA对)。Spatial-DISE通过结合现实世界数据收集和Blender软件的合成生成,旨在解决现有基准在评估空间推理能力方面的不足,特别是内在动态空间推理。数据集提供了一个结构化和可重复的方法来理解模型的失败,并通过可扩展和可验证的数据生成流程,为细粒度评估和未来模型开发提供了宝贵资源。
Spatial-DISE is a unified benchmark for evaluating the spatial reasoning capabilities of Vision-Language Models (VLMs). Developed by a research team from the Department of Computer Science, University of Liverpool, the dataset comprises over 12K validated spatial reasoning Visual Question Answering (VQA) pairs. It is split into two subsets: Spatial-DISE Bench (559 evaluation VQA pairs) and Spatial-DISE-12K (12K+ training VQA pairs). Constructed by combining real-world data collection and synthetic generation using Blender software, Spatial-DISE aims to address the shortcomings of existing benchmarks in evaluating spatial reasoning capabilities, particularly intrinsic dynamic spatial reasoning. The dataset provides a structured and reproducible framework for understanding model failures, and offers valuable resources for fine-grained evaluation and future model development via a scalable and verifiable data generation pipeline.




