UniETP
收藏资源简介:
UniETP是由北京大学研究团队构建的具身任务规划统一基准数据集,旨在解决现有模拟器与数据集碎片化问题,推动通用化具身智能体发展。该数据集整合了AI2-THOR、VirtualHome、Habitat和BEHAVIOR四大主流模拟器,通过自动化任务生成流程创建了超过500个任务实例,涵盖138种任务类型,并支持多粒度观察与动作空间。数据集构建过程基于100余个任务模板,结合常识知识库与大语言模型标注器,系统化生成具有不同逻辑复杂度、实例级 grounding 和语言理解难度的任务。其核心应用于评估具身任务规划模型的泛化能力,特别是在跨模拟器环境下的指令理解、复杂逻辑推理和实例级交互规划等关键挑战。
UniETP is a unified benchmark dataset for embodied task planning constructed by a research team at Peking University, which aims to address the fragmentation issue of existing simulators and datasets and advance the development of general-purpose embodied AI agents. This dataset integrates four mainstream simulators including AI2-THOR, VirtualHome, Habitat and BEHAVIOR, generates over 500 task instances through an automated task generation pipeline, covers 138 task types, and supports multi-granularity observation and action spaces. The dataset construction process is based on more than 100 task templates, combined with common sense knowledge bases and large language model-based annotators, to systematically produce tasks with varying logical complexity, instance-level grounding and language understanding difficulty. Its core application is to evaluate the generalization capability of embodied task planning models, particularly regarding key challenges such as instruction understanding, complex logical reasoning and instance-level interactive planning across simulator environments.
UniETP 数据集概述
数据集名称: UniETP(Unifying Environments for Generalizable Embodied Task Planning)
核心定位: 用于开发和评估可泛化具身任务规划智能体的统一基准。
主要特点
- 四合一模拟器接口: 集成 AI2-THOR、VirtualHome、Habitat 和 BEHAVIOR 四种模拟器,提供统一的观察空间、动作空间、场景图表示和评估协议。
- 可调节评估难度: 提供三种模式(M1/M2/M3),分别调整场景先验、动作粒度和定位要求。
- 丰富的任务覆盖: 包含 138 个任务模板,覆盖对象状态、空间关系、时间约束、量化计数和组合目标。
- 自动任务生成: 支持任务模板自动实例化、场景落地、脚本专家策略验证,并转化为自然语言指令。
- 细粒度诊断分析: 按能力维度划分任务,暴露长程规划、场景探索和实例级视觉定位的瓶颈。
评估能力维度
- 原子能力
- 组合能力
- 逻辑推理
- 语言理解
- 实例定位
开发状态
数据与代码正在积极开发中,敬请期待后续更新。

- 1UniETP: Unifying Environments for Generalizable Embodied Task Planning北京大学 · 2026年



