HABIT
收藏资源简介:
HABIT(Human-Aware Behavior and Interaction Training dataset)是一个专为人类在场环境设计的大规模机器人演示数据集,旨在解决现有数据集在人类缺席环境中收集导致策略难以在真实部署环境(如家庭、工厂)中与人类协调的问题。它包含10,563个演示片段,总计164.19小时的双手机器人操作数据,覆盖60个不同任务。任务根据人机交互依赖关系分为三种角色:协作角色(人机共同完成共享目标,涉及直接物理交互)、同事角色(共享目标和空间但独立执行任务,需避免碰撞)和监管角色(人类通过手势等指令指导机器人)。数据收集采用双操作员模式,使用双Franka Research 3机械臂和Robotiq 2F-85夹爪,通过五个RGB摄像头流(三个机器人侧视角和两个人类侧视角)记录,动作数据以关节空间、笛卡尔空间和夹爪状态多种形式记录。数据集提供精细的子任务级标注,实时记录人类和机器人子任务边界,实现时间步精确对齐。采用LeRobot v2.0格式组织,包含元数据(任务指令、子任务指令、统计信息等)、Parquet格式的状态-动作序列文件和MP4格式视频文件,适用于训练视觉-语言-动作模型和世界动作模型,以提升机器人在人机共存环境中的任务成功率和社会兼容性行为。
HABIT (Human-Aware Behavior and Interaction Training dataset) is a large-scale robotic demonstration dataset specifically designed for human-present environments, aiming to address the issue that existing datasets collected in human-absent scenarios make it difficult for robot policies to coordinate with humans in real-world deployment environments such as homes and factories. It contains 10,563 demonstration segments, totaling 164.19 hours of dual-arm robotic manipulation data, covering 60 distinct tasks. Tasks are categorized into three roles based on human-robot interaction dependencies: collaborative role (human and robot jointly complete shared goals involving direct physical interaction), coworker role (share goals and space but perform tasks independently, requiring collision avoidance), and supervisory role (humans instruct robots via gestures and other commands). The data was collected in a dual-operator setup, using dual Franka Research 3 robotic arms and Robotiq 2F-85 grippers. Data was recorded via five RGB camera streams (three robot-side viewpoints and two human-side viewpoints), and motion data was captured in multiple forms including joint space, Cartesian space, and gripper status. The dataset provides fine-grained subtask-level annotations, which record the boundaries of human and robot subtasks in real time and enable precise time-step alignment. Organized in the LeRobot v2.0 format, the dataset includes metadata (task instructions, subtask instructions, statistical information, etc.), Parquet-formatted state-action sequence files and MP4-format video files. It is suitable for training vision-language-action models and world action models to improve the task success rate and socially compliant behaviors of robots in human-robot coexistence environments.
HABIT 数据集概述
HABIT(Human-Aware Behavior and Interaction Training Dataset)是一个面向人类在场环境的机器人操作演示数据集,旨在训练机器人策略学习人类感知行为。
- 许可协议: CC-BY-4.0
- 任务类别: 机器人学
- 语言: 英语
- 关键词: 机器人操作数据集、人机交互、视觉-语言-动作模型
- 规模: 10,000 < N < 100,000 个样本
核心特点
- 人类在场演示: 每个数据片段都包含一位在场的人类搭档,在响应式收集协议下采集。
- 三种交互角色: 合作者 (Collaborator)、同事 (Coworker)、监督者 (Supervisor),三种角色的片段数量相当。
- 任务工作流: 采用图结构任务表示,明确捕获跨智能体的依赖关系。
- 子任务级别标注: 在演示过程中实时记录机器人和人类的子任务边界。
- 五路RGB摄像流: 3个机器人侧视角 + 2个人类侧视角,全面捕捉人机交互。
- 双臂操作: 使用两台Franka Research 3 (FR3) 机械臂。
数据集统计
- 任务数量: 60
- 数据片段数: 10,563
- 总帧数: 591万
- 总时长: 164.19 小时
- 交互角色: 3种(合作者/同事/监督者)
- 每个片段摄像头数量: 5个
- 机器人平台: 双臂 Franka Research 3
- 数据格式: LeRobot v2.0
片段时长(按任务统计): 均值 59.9 秒,中位数 56.4 秒,范围 30.3 – 101.4 秒。
独特子任务: 157个机器人子任务,182个人类子任务,308个独特的人-机器人子任务对。
按角色细分统计
| 角色 | 任务数 | 片段数 | 时长(小时) | 机器人子任务/片段 | 人类子任务/片段 |
|---|---|---|---|---|---|
| 合作者 | 20 | 3,198 | 49.98 | 3.86 | 4.70 |
| 同事 | 20 | 3,969 | 57.33 | 3.12 | 4.25 |
| 监督者 | 20 | 3,396 | 56.88 | 4.15 | 3.98 |
三种交互角色详解
- 合作者: 人和机器通过直接物理交互(如递物、共提水桶)共同完成一个共享目标。机器人必须在空间和时间上与人类协调。
- 同事: 人和机器共享一个目标和空间,但无直接物理接触。每个智能体独立处理自己的任务部分。机器人必须避免与人类碰撞以确保安全。
- 监督者: 人类通过明确线索(如手势或行为指令)指挥机器人,机器人必须仅从视觉输入感知人类的意图。
硬件设置
- 机器人: 双臂 Franka Research 3,配备 Robotiq 2F-85 夹爪。
- 工作空间: 一个前桌(人类与机器人之间,直接共享工作空间)和一个侧桌(人类一侧,用于人类活动)。
- 摄像头(5个RGB):
- 1个机器人中心摄像头:向前倾斜,捕获人类和共享工作空间。
- 2个腕部摄像头:每臂一个。
- 1个人类头戴式摄像头(第一人称视角)。
- 1个外部摄像头:观察整个工作空间。
- 遥操作: 使用 Meta Quest 3 控制器,基于 DROID 代码库。动作表示以关节空间、笛卡尔空间和夹爪状态形式记录。
数据收集协议
- 每个片段由两位操作员记录:机器人操作员(遥控双臂 FR3)和人类操作员(执行人类侧活动)。
- 每位操作员通过脚踩踏板在线标记子任务边界。
- 协议三原则:
- 响应式交互: 每位操作员仅在直接观察搭档后行动,禁止排练、眼神信号或口头指令等场外协调。
- 目标行为引导: 通过特定设计引导数据中出现如让行、时间适应和手势理解等人类感知行为。
- 多样性: 在片段间变化服装颜色和物体交互顺序,并涉及多位不同体型的人类操作员。
任务详情
数据集包含60个任务,涵盖三种角色。完整的任务列表(高级指令)位于 meta/tasks.jsonl 中。每个片段在 meta/episodes.jsonl 中通过 sid(场景ID)字段链接到详细的 task_details/ 目录下的文档,涵盖环境设置、高级指令、低级指令和工作流图。
数据集结构
数据遵循 LeRobot v2.0 格式。 full 和 sample 两个配置独立组织。目录结构包括 meta/(元数据)、data/(Parquet文件存储完整片段的状态、动作、时间戳和语言指令)、videos/(MP4视频文件)。
子任务标注: 每个步骤(step)的 Parquet 文件中都携带当前活动的子任务索引:
low_level_task_index: 机器人子任务索引。human_role_subtask_index: 人类子任务索引。
摄像头特征键:
| 流 | 侧 | 用途 |
|---|---|---|
observation.images.front_view |
机器人 | 正前方中心视图;捕获人类和共享工作区域 |
observation.images.left_wrist_view |
机器人 | 左臂腕部摄像头 |
observation.images.right_wrist_view |
机器人 | 右臂腕部摄像头 |
observation.images.human_front_view |
人类 | 头戴式,人类第一人称视角 |
observation.images.exo_view |
外部 | 全视角外部摄像头 |
引用
bibtex @article{song2026habit, author = {Jaehwi Song and Suchae Jeong and Byeongguk Jeon and Sungdong Kim and Minjoon Seo and Hyungmok Son and Kimin Lee}, title = {HABIT: Human-Aware Behavior and Interaction Training Dataset for Robot Manipulation}, journal = {arXiv preprint arXiv:2606.31682}, year = {2026} }




