HyeonseokE/SO101-cap_pnp_baseline_10fps_40epi
收藏资源简介:
该数据集是一个机器人数据集,由LeRobot创建,专注于机器人任务。数据集采用so101_follower机器人类型,包含40个总片段、12463个总帧和1个总任务。数据以10帧每秒的速率采集,并分割为训练集(覆盖所有片段)。数据集结构包括多个特征:观察状态(如关节位置,包括肩部平移、肩部升降、肘部弯曲、腕部弯曲、腕部旋转和夹爪位置,形状为6维浮点数组)、动作(与观察状态类似,形状为6维浮点数组)、观察图像(包括顶部和左腕视角的视频数据,分辨率为480x640,3通道,使用av1编解码器,帧率为30)、末端执行器位置(机器人xyzrpy坐标,形状为6维浮点数组)、夹爪二进制状态(1维浮点数组)、技能信息(如自然语言描述、验证问题、类型、进度、目标位置关节坐标、目标位置机器人xyzrpy坐标、目标位置夹爪状态,其中自然语言和验证问题为字符串类型,其他为浮点或整数类型)、子任务信息(如自然语言描述、对象名称、目标位置坐标,形状为3维浮点数组)、时间戳、帧索引、片段索引、索引和任务索引(均为整数类型)。数据以parquet文件格式存储,总数据文件大小为100MB,视频文件大小为200MB,分块大小为1000。数据集适用于机器人学习和控制任务,提供丰富的多模态数据支持。
This dataset is a robotics dataset created using LeRobot, focusing on robot tasks. It employs the so101_follower robot type, containing 40 total episodes, 12463 total frames, and 1 total task. Data is collected at 10 frames per second and split into a training set (covering all episodes). The dataset structure includes multiple features: observation state (e.g., joint positions including shoulder pan, shoulder lift, elbow flex, wrist flex, wrist roll, and gripper positions, shaped as a 6-dimensional float array), action (similar to observation state, shaped as a 6-dimensional float array), observation images (including video data from top and left wrist perspectives, with a resolution of 480x640, 3 channels, using av1 codec, at 30 fps), end-effector position (robot xyzrpy coordinates, shaped as a 6-dimensional float array), gripper binary state (1-dimensional float array), skill information (such as natural language description, verification question, type, progress, goal position joint coordinates, goal position robot xyzrpy coordinates, goal position gripper state, with natural language and verification questions as string types, others as float or integer types), subtask information (e.g., natural language description, object name, target position coordinates, shaped as a 3-dimensional float array), timestamp, frame index, episode index, index, and task index (all integer types). Data is stored in parquet file format, with a total data file size of 100MB, video file size of 200MB, and chunk size of 1000. The dataset is suitable for robot learning and control tasks, providing rich multimodal data support.




