Hy-Embodied-0.5-VLA-Data
收藏资源简介:
Hy-Embodied-0.5-VLA-Data 是由腾讯 Robotics X 与腾讯混元团队发布的大规模双臂机器人操作数据集,用于训练视觉-语言-动作(Vision-Language-Action, VLA)基础模型。该数据集基于定制化指尖式 UMI(Universal Manipulation Interface)设备结合光学动作捕捉系统采集,包含超过 2,000 小时的高保真机器人操作演示数据,覆盖 70 余类双臂灵巧操作任务。公开版本包含约 250,304 个操作轨迹(episodes)、233,600,314 帧数据,总规模约 18.8 TB,数据以 Lance 格式存储,并兼容 LeRobot v3.0 数据规范。每条数据记录包含多视角 RGB 图像(头部相机、左右腕部相机)、双臂末端执行器状态、夹爪状态、动作信息以及任务文本描述,可支持机器人模仿学习、视觉语言动作模型预训练、多任务操作学习以及跨机器人平台迁移研究。
Hy-Embodied-0.5-VLA-Data is a large-scale dual-arm robotic manipulation dataset released by Tencent Robotics X and Tencent Hunyuan Team, designed for training Vision-Language-Action (VLA) foundation models. This dataset is collected using a customized fingertip-type UMI (Universal Manipulation Interface) device combined with an optical motion capture system, containing over 2,000 hours of high-fidelity robotic manipulation demonstration data covering more than 70 categories of dual-arm dexterous manipulation tasks. The public version includes approximately 250,304 manipulation episodes, 233,600,314 frames of data, with a total size of around 18.8 TB. The data is stored in Lance format and compatible with the LeRobot v3.0 data specification. Each data record contains multi-view RGB images (head camera, left and right wrist cameras), dual-arm end-effector states, gripper states, action information, and task text descriptions, which can support research on robotic imitation learning, visual-language-action model pre-training, multi-task manipulation learning, and cross-robotic-platform transfer learning.




