MSMU
收藏资源简介:
MSMU数据集是一个大规模定量空间推理数据集,包含约25K图像和700K问答对(包括10K思维链样本),来自2K真实3D场景,带有2.5M数值注释。该数据集旨在提升视觉语言模型的空间感知能力,特别是在定量空间推理方面。数据集的创建过程包括场景图构建、3D到2D映射、问答生成等步骤。MSMU数据集的应用领域包括机器人、自动驾驶汽车、增强现实等,旨在解决现有视觉语言模型在理解3D空间概念方面的不足。
The MSMU dataset is a large-scale quantitative spatial reasoning dataset, containing approximately 25K images and 700K question-answer pairs (including 10K chain-of-thought samples) sourced from 2K real 3D scenes, with 2.5M numerical annotations. This dataset aims to enhance the spatial perception capabilities of vision-language models, especially in the domain of quantitative spatial reasoning. The dataset construction process includes steps such as scene graph construction, 3D-to-2D mapping, and question-answer generation. The application areas of the MSMU dataset cover robotics, autonomous driving, augmented reality and other fields, and it is designed to address the shortcomings of existing vision-language models in understanding 3D spatial concepts.




