MultiSPA
收藏资源简介:
MultiSPA是一个大规模的多帧空间理解数据集,包含超过2700万个样本,涵盖了多样化的3D和4D场景。该数据集支持多种模态的引用和输出格式,包括视觉点注释、像素坐标和语义标签,从而拓宽了潜在的应用场景。数据集包含了从文本到标量、二维像素位置和三维位移向量等多种类型的空间信息。研究人员利用现有的注释3D和4D数据集进行数据收集,并通过采样具有均匀重叠分布的图像对以及回投影空间和时间对齐的点云来建立像素对应关系。MultiSPA数据集旨在帮助多模态大型语言模型更好地理解多帧空间信息,并用于机器人等实际应用中的空间推理任务。
MultiSPA is a large-scale multi-frame spatial understanding dataset containing over 27 million samples and covering diverse 3D and 4D scenarios. This dataset supports multiple modal input and output formats, including visual point annotations, pixel coordinates, and semantic labels, thus expanding its potential application scenarios. The dataset encompasses various types of spatial information ranging from text to scalars, 2D pixel positions, and 3D displacement vectors. Researchers collected data using existing annotated 3D and 4D datasets, and established pixel correspondences by sampling image pairs with uniformly overlapping distributions and back-projecting spatially and temporally aligned point clouds. The MultiSPA dataset aims to assist multimodal large language models in better understanding multi-frame spatial information, and supports spatial reasoning tasks in real-world applications such as robotics.
Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models
概述
Multi-SpatialMLLM 是一个多模态大语言模型(MLLM),专注于多帧空间理解。它能够感知相机和物体运动、深度、视觉对应关系、物体大小等,支持不同类型的输入引用和输出。
数据集与基准
- MultiSPA数据集:包含超过2700万个样本,涵盖多样化的3D和4D场景。
- MultiSPA基准:专注于多帧空间理解,采用统一的准确性指标,每个任务有各自的真阳性结果判定标准。
数据引擎与生成
- 静态数据引擎:为每个静态场景根据重叠率采样图像对,并计算元空间信息以构建问答对。
- 刚体分割:使用4D数据集(4D物体跟踪数据集)构建物体运动感知数据,通过刚体分割方法确保多样性。
- 数据生成模块:更多细节见论文。
实验成果
- MultiSPA基准表现:Multi-SpatialMLLM在定性和定量子任务上显著优于基线,平均提升36%,超越更大的专有模型。
- 泛化性能:在BLINK数据集和通用VQA基准上表现出强大的泛化能力。
- 可扩展性能:通过增加可训练参数和训练数据,性能进一步提升。
- 涌现能力:在困难的空间理解任务(如视觉对应任务)中,仅大型模型通过微调能取得更好表现。
机器人应用
- 可作为多帧奖励标注器,用于机器人学习。
- 能够预测给定两帧中目标物体的移动距离。
参考文献
bibtex @article{xu2025multi, title={Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models}, author={Xu, Runsen and Wang, Weiyao and Tang, Hao and Chen, Xingyu and Wang, Xiaodong and Chu, Fu-Jen and Lin, Dahua and Feiszli, Matt and Liang, Kevin J.}, journal={arXiv preprint arXiv:2505.17015}, year={2025} }

- 1Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language Models香港中文大学 · 2025年



