遇见数据集

PICK-CORD: Transferring Human Manipulation to Robots via Multi-Sensor Human Demonstrated Dataset for Sequential Action Planning

收藏
Zenodo2026-09-26 更新2026-10-01 收录
官方服务:

资源简介:

The increasing integration of robots into human-centric environments requires efficient methodsfor transferring task knowledge from humans to machines. Learning from Demonstration (LfD) offersa promising paradigm for this transfer, yet existing human-demonstrated datasets often lack the multimodalrichness needed for robust trajectory learning, typically capturing only partial information such as RGBDimagery or kinematic joint data in isolation. This paper presents PICK-CORD, a human-demonstrateddataset for pick-and-place task planning that synchronously integrates RGB-D video, Inertial MeasurementUnit (IMU) hand orientation data, and natural language task captions across 119 recorded trials. Eachsample is structured into an Approaching Segment and a Task Segment, jointly forming a complete pickand place "Action," and is delivered with full per-frame timestamps to support precise temporal alignment.A preprocessing pipeline is described that performs depth frame interpolation, vision-based hand detection,marker-aided camera calibration, and inertial data downsampling to extract synchronized 3D trajectories andorientation profiles from the raw sensor streams. To validate the dataset’s utility, a baseline convolutionalneural network was trained to predict five-point manipulation trajectories and end-effector orientation frominitial scene images, depth maps, and object masks. In offline evaluation, the baseline model achieved a100% success rate under a 3 cm path accuracy threshold, with orientation accuracy of 100% for roll andpitch and 90.91% for yaw. When deployed on a physical Delta parallel robot across nine real-world pick-and-place tasks spanning three difficulty levels, the system achieved a 77.77% functional success rate, includingtasks involving a previously unseen object. These results validate the dataset’s high efficiency and utility bydemonstrating that its visual and inertial modalities successfully supervise robust trajectory and orientationlearning, while the synchronized linguistic annotations provide a rich foundation for future vision-languageresearch.

提供机构:
Zenodo
创建时间:
2026-09-26
二维码
社区交流群
二维码
科研交流群
商业服务