Perspective Taking; Path Tracing; Multiview Counting
收藏资源简介:
该研究构建了三个空间想象力数据集,由华盛顿大学、艾伦人工智能研究所等机构联合创建,旨在训练和评估视觉语言模型的空间推理能力。数据集包含约8.6万条样本,涵盖模拟和真实环境,数据来源于AI2-THOR、Habitat、Matterport3D等平台,通过渲染和模板生成技术创建。这些数据集应用于空间推理任务,解决视角转换、路径追踪和多视图计数等问题,提升模型对未观察空间结构的想象力。
This study constructs three spatial imagination datasets, jointly created by institutions including the University of Washington, the Allen Institute for AI, and others. The datasets are designed to train and evaluate the spatial reasoning capabilities of vision-language models. They contain approximately 86,000 samples, covering both simulated and real-world environments, with data sourced from platforms such as AI2-THOR, Habitat, and Matterport3D, and generated through rendering and template generation techniques. These datasets are applied to spatial reasoning tasks, addressing problems like viewpoint transformation, path tracing, and multi-view counting, to enhance the model's imagination of unobserved spatial structures.




