MansionWorld
收藏资源简介:
MansionWorld 是一个面向具身智能的建筑级多楼层三维环境数据集,基于语言驱动的 MANSION 框架生成,旨在支持跨楼层导航与长时序任务的研究。该数据集构建了包含住宅、办公楼及公共设施在内的多类型三维建筑场景,每栋建筑具有 2–10 层结构,并通过楼梯、电梯等实现真实的垂直连通关系,从而模拟复杂的室内空间拓扑。数据集提供结构化的场景布局(如 JSON 格式平面结构)与对应的视觉表示(如平面图与可交互环境),并结合语义信息与任务描述,用于刻画智能体在多楼层环境中的感知、决策与执行过程。MansionWorld 包含上千个可交互建筑与上万个功能空间,适用于跨楼层导航、目标搜索、任务规划、多步操作及空间推理等任务研究。该数据集特别适合具身智能、机器人导航、3D 场景理解与强化学习等方向,可用于训练视觉-语言-行动(VLA)模型、具身大模型及长时序决策模型,推动智能体从“房间级”向“建筑级”复杂环境能力的提升。
MansionWorld is a building-level, multi-floor 3D environment dataset tailored for embodied intelligence, generated using the language-driven MANSION framework. It is designed to support research on cross-floor navigation and long-horizon tasks. This dataset constructs diverse 3D architectural scenes covering residential buildings, office buildings and public facilities. Each building features a 2–10 floor structure, with realistic vertical connectivity implemented via stairs, elevators and other means, thereby simulating complex indoor spatial topologies. The dataset provides structured scene layouts (e.g., JSON-formatted planar structural data) and corresponding visual representations (e.g., floor plans and interactive environments), paired with semantic information and task descriptions to characterize the perception, decision-making and execution processes of AI agents in multi-floor environments. MansionWorld includes thousands of interactive buildings and tens of thousands of functional spaces, making it applicable to research on tasks such as cross-floor navigation, target search, task planning, multi-step operation and spatial reasoning. This dataset is particularly well-suited for research directions including embodied intelligence, robotic navigation, 3D scene understanding and reinforcement learning, and can be used to train vision-language-action (VLA) models, embodied large language models (LLMs) and long-horizon decision-making models, advancing the improvement of intelligent agents' capabilities from "room-level" to "building-level" complex environments.




