SpatialForge
收藏资源简介:
SpatialForge-10M是一个专为从开放世界2D图像中进行3D感知的空间感知与推理而设计的大规模视觉-语言数据集。它包含超过1000万个问答对,源自280万张经过精心筛选的真实世界图像,覆盖从低层空间感知到高层空间推理的多种任务。图像数据来源于多样化的开放世界图像集(如Objects365、Pixmo和OpenImages),确保了在室内、室外、自我中心视角及互联网规模场景中的广泛视觉多样性。数据集旨在支持空间推理预训练、多任务视觉语言模型监督、3D感知学习、基础与指代研究、相机中心与人类中心推理以及空间指令微调。核心内容包含六大空间任务,分为感知任务(包括基础定位、指代描述和对象计数)和关系推理任务(包括远近深度判断、左右空间关系推断以及人类中心视角推理)。数据以统一的问答对格式组织,适用于训练现代多模态大语言模型(如Qwen-VL、InternVL、LLaVA),关键特征包括覆盖对象中心与人类中心的推理、边界框坐标归一化为[0, 1000]的格式(遵循Qwen3-VL预训练惯例),并强调从2D图像中构建几何一致且具有视角感知的空间监督信号。
SpatialForge-10M is a large-scale vision-language dataset designed for spatial perception and reasoning with 3D awareness from open-world 2D images. It contains over 10 million question-answer pairs derived from 2.8 million carefully curated real-world images, covering a range of tasks from low-level spatial perception to high-level spatial reasoning. The image data is sourced from diverse open-world image collections, including Objects365, Pixmo, and OpenImages, ensuring broad visual diversity across indoor, outdoor, egocentric, and internet-scale scenes. The dataset aims to support spatial reasoning pre-training, multi-task vision-language model supervision, 3D-aware learning, grounding and referring research, camera-centric and human-centric reasoning, and spatial instruction tuning. Its core content comprises six spatial tasks, divided into two categories: perception tasks (including grounding localization, referring description, and object counting) and relational reasoning tasks (including far-near depth judgment, left-right spatial relationship inference, and human-centric perspective reasoning). Data is organized in a unified question-answer pair format, suitable for training modern multimodal large language models such as Qwen-VL, InternVL, and LLaVA. Key features include coverage of object-centric and human-centric reasoning, bounding box coordinates normalized to [0, 1000] (following Qwen3-VL pre-training conventions), and an emphasis on constructing geometrically consistent and viewpoint-aware spatial supervision signals from 2D images.
数据集概述:SpatialForge-10M
SpatialForge-10M 是一个大规模视觉语言数据集,专为从开放世界二维图像中学习三维感知的空间感知与推理而设计。
核心数据
- 规模:包含超过 1000 万 个问答对,源自 280 万 张经过筛选的真实世界图像。
- 许可:Creative Commons Attribution Non Commercial 4.0 (cc-by-nc-4.0)。
- 语言:英语。
- 数据来源:基于大规模的公开图像数据集构建,包括 Objects365、Pixmo 和 OpenImages,覆盖室内、室外、自我中心视角及互联网规模的多样化场景。
任务架构
数据集包含六大空间任务,分为感知和关系推理两个层级:
| 层级 | 任务 | 描述 | 数量 |
|---|---|---|---|
| 感知 | 定位 (Grounding) | 根据文本描述定位物体 → 预测边界框 | 360 万 |
| 指代 (Referring) | 根据区域/边界框生成物体描述 | 360 万 | |
| 计数 (Counting) | 计算满足语义条件的物体数量 | 49.5 万 | |
| 关系推理 | 远近 (Near-Far) | 判断物体之间的相对深度 | 260 万 |
| 左右 (Left-Right) | 推断以相机为中心的水平空间关系 | 9.3 万 | |
| 视角 (Perspective) | 以人为中心的视角与空间推理 | 0.8 万 | |
| 总计 | 1020 万 |
关键特性
- 统一格式:采用统一的问答格式,适用于训练现代多模态大语言模型,如 Qwen-VL、InternVL 和 LLaVA。
- 边界框格式:边界框坐标归一化至 [0, 1000] 范围内。
- 设计目标:适用于空间推理预训练、多任务视觉语言模型监督、三维感知学习、指代与定位研究、以相机/人为中心的推理以及空间指令微调。
重要说明
该数据集仅提供完整的注释文件(问答对与任务划分),用户需要自行从图像来源数据集(Objects365、OpenImages、Pixmo)获取对应的原始图像。




