AxonData/human-movement-pov-dataset-for-robotics
收藏资源简介:
VLA第一人称视角视频数据集——100+小时的4K人类操作视频,包含100多个小时的真人第一人称视角(POV)视频,记录了真实人类执行日常生活手部任务的过程,用于训练视觉-语言-动作(VLA)模型、模仿学习策略和具身AI系统。视频为连续未剪辑的4K分辨率(30 FPS),涵盖维修、组装、缝纫、家务、户外工作、电子设备、园艺、自行车维护等多种任务,采用头戴式视角以匹配机器人传感器几何结构,并经过人工审核确保手部可见性和任务连续性。该数据集提供商业许可,适用于生产级机器学习训练、VLA预训练和基础模型训练。
VLA Egocentric Video Dataset — 100+ Hours of 4K Human Manipulation, featuring 100+ hours of first-person (egocentric POV) video of real humans performing real-life hand tasks for training vision-language-action (VLA) models, imitation learning policies, and embodied AI systems. The videos are continuous, uncut 4K resolution (30 FPS) footage capturing full task arcs, recorded from head-mounted perspectives to better match robot sensor geometry, and cover diverse tasks such as repair, assembly, sewing, household chores, outdoor work, electronics, gardening, and bike maintenance. Each clip is manually verified for hand visibility and task continuity, and the dataset is available with a commercial license for production ML training, VLA pretraining, and foundation model training.




