遇见数据集

Hy-Embodied-0.5-VLA-Data

收藏
魔搭社区2026-07-11 更新2026-07-05 收录
官方服务:

资源简介:

Hy-Embodied-0.5-VLA-Data 是由腾讯 Robotics X 与腾讯混元团队发布的大规模双臂机器人操作数据集,用于训练视觉-语言-动作(Vision-Language-Action, VLA)基础模型。该数据集基于定制化指尖式 UMI(Universal Manipulation Interface)设备结合光学动作捕捉系统采集,包含超过 2,000 小时的高保真机器人操作演示数据,覆盖 70 余类双臂灵巧操作任务。公开版本包含约 250,304 个操作轨迹(episodes)、233,600,314 帧数据,总规模约 18.8 TB,数据以 Lance 格式存储,并兼容 LeRobot v3.0 数据规范。每条数据记录包含多视角 RGB 图像(头部相机、左右腕部相机)、双臂末端执行器状态、夹爪状态、动作信息以及任务文本描述,可支持机器人模仿学习、视觉语言动作模型预训练、多任务操作学习以及跨机器人平台迁移研究。

Hy-Embodied-0.5-VLA-Data is a large-scale dual-arm robotic manipulation dataset released by Tencent Robotics X and Tencent Hunyuan Team, designed for training Vision-Language-Action (VLA) foundation models. This dataset is collected using a customized fingertip-type UMI (Universal Manipulation Interface) device combined with an optical motion capture system, containing over 2,000 hours of high-fidelity robotic manipulation demonstration data covering more than 70 categories of dual-arm dexterous manipulation tasks. The public version includes approximately 250,304 manipulation episodes, 233,600,314 frames of data, with a total size of around 18.8 TB. The data is stored in Lance format and compatible with the LeRobot v3.0 data specification. Each data record contains multi-view RGB images (head camera, left and right wrist cameras), dual-arm end-effector states, gripper states, action information, and task text descriptions, which can support research on robotic imitation learning, visual-language-action model pre-training, multi-task manipulation learning, and cross-robotic-platform transfer learning.

提供机构:
maas
创建时间:
2026-06-15
搜集汇总
数据集介绍
Hy-Embodied-0.5-VLA-Data 数据集图片
背景与挑战
背景概述
Hy-Embodied-0.5-VLA-Data是一个大规模双手操作数据集,专为训练视觉-语言-动作(VLA)基础模型设计,基于光学运动捕捉通过自定义指尖UMI设备收集了2000+小时的高保真演示,覆盖70多个操作任务。数据集以Lance格式发布(兼容LeRobot v3.0),包含约2,163小时的总时长、250,304个episodes和约18.8 TB的数据,支持多摄像头视图(头部和双腕摄像头),开源版本约占完整语料库的20%。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务