shana643/SpatialForge
收藏资源简介:
SpatialForge-10M是一个大规模视觉语言数据集,专为从开放世界2D图像进行3D感知空间感知和推理而设计。该数据集包含超过1000万个问答对,这些问答对源自280万张精选的真实世界图像,覆盖了低层空间感知和高层空间推理任务。数据集构建自多样化的开放世界图像源,包括Objects365、Pixmo和OpenImages,确保了室内、室外、自我中心和互联网规模场景的广泛视觉多样性。它适用于空间推理预训练、多任务视觉语言模型监督、3D感知学习、接地和指代研究、相机中心和人本中心推理以及空间指令调优,并支持统一问答格式,适合训练现代多模态大语言模型,如Qwen-VL、InternVL和LLaVA。
SpatialForge-10M is a large-scale vision-language dataset designed for 3D-aware spatial perception and reasoning from open-world 2D images. The dataset contains over 10 million QA pairs generated from 2.8 million curated real-world images, covering both low-level spatial perception and high-level spatial reasoning tasks. It is constructed from diverse open-world image sources including Objects365, Pixmo and OpenImages, enabling broad visual diversity across indoor, outdoor, egocentric, and internet-scale scenes. The dataset is designed for spatial reasoning pretraining, multitask VLM supervision, 3D-aware perception learning, grounding and referring research, camera-centric and human-centric reasoning, and spatial instruction tuning, and supports a unified QA format suitable for training modern multimodal large language models such as Qwen-VL, InternVL, and LLaVA.




