dvla-can2k-futee-sweep
收藏资源简介:
该数据集包含四个LeRobot数据集,基于相同的2005个MuJoCo/robosuite演示。所有四个变体的图像、状态和片段边界完全相同,仅动作标签不同,差异在于未来末端执行器重新标记窗口的大小(make_lerobot --futee_offset)。任务为“放置”:从桌子上拿起一个罐头并放入碗中。80%的片段开始时罐头处于滚动状态(速度均匀分布在0.25–1.5 m/s),其余为静态。演示来自具有特权模拟器状态的脚本状态机,仅保留成功尝试。单个物体(can11),五个容器。控制频率250 Hz,Panda臂在差分IK和刚性关节位置控制器下运行。数据集布局为LeRobot v2.1,robot_type: panda,3个RGB摄像头(360×480),以及每个摄像头的v2e事件极性渲染。动作(action)为10维向量,为delta形式:绝对目标位姿减去同一帧的状态。评估时,策略输出解码为target = live_state + predicted_delta,每一步重新锚定。四个变体:futee0(偏移0,前瞻0ms,帧数1,695,122),futee20(偏移20,前瞻80ms,帧数1,705,362),futee40(偏移40,前瞻160ms,帧数1,745,478),futee60(偏移60,前瞻240ms,帧数1,785,578)。推荐使用futee0变体。
This dataset contains four LeRobot datasets, based on the same 2005 MuJoCo/robosuite demonstrations. All four variants share identical images, states, and episode boundaries, differing only in action labels due to variations in the future end-effector re-labeling window size (make_lerobot --futee_offset). The task is place: pick up a can from the table and place it into a bowl. 80% of episodes start with the can rolling (velocity uniformly distributed between 0.25–1.5 m/s), while the remaining are static. Demonstrations are from a scripted state machine with privileged simulator state, retaining only successful attempts. Single object (can11), five containers. Control frequency is 250 Hz, Panda arm operates under differential IK and rigid joint position controller. Dataset layout is LeRobot v2.1, robot_type: panda, 3 RGB cameras (360×480), and v2e event polarity rendering for each camera. Action is a 10-dimensional vector in delta form: absolute target pose minus the state at the same frame. During evaluation, policy output is decoded as target = live_state + predicted_delta, re-anchored at each step. Four variants: futee0 (offset 0, lookahead 0ms, frames 1,695,122), futee20 (offset 20, lookahead 80ms, frames 1,705,362), futee40 (offset 40, lookahead 160ms, frames 1,745,478), futee60 (offset 60, lookahead 240ms, frames 1,785,578). The futee0 variant is recommended.
数据集概述:dvla-can2k-futee-sweep
基本信息
- 许可证:Apache-2.0
- 任务类别:机器人(robotics)
- 标签:LeRobot、MuJoCo、robosuite、Panda、操作(manipulation)、动态抓取(dynamic-grasping)
- 数据集格式:LeRobot v2.1,包含4个子配置(
futee0、futee20、futee40、futee60),每个配置对应不同的动作标注前瞻偏移量。
核心内容
该数据集基于同一组2,005条MuJoCo/robosuite演示构建,四个变体共享相同的图像、状态和回合边界,唯一区别在于动作标签的标注方式——通过--futee_offset参数控制未来末端执行器重标注窗口的大小。
四个变体对比
| 变体 | 偏移量 | 前瞻时间 @250Hz | 总帧数 | 裁剪的起始保持帧数 |
|---|---|---|---|---|
futee0 |
0 | 0 ms | 1,695,122 | 55.0 |
futee20 |
20 | 80 ms | 1,705,362 | 49.9 |
futee40 |
40 | 160 ms | 1,745,478 | 29.9 |
futee60 |
60 | 240 ms | 1,785,578 | 9.9 |
另有一个320 ms变体(futee80)存放在独立仓库中:dvla-can2k-fixed-250hz-events-250fps-delta-trim。
任务描述
任务名称:place(放置)
- 将罐头从桌面上拾起并放入碗中。
- 80% 的回合开始时罐头处于滚动状态(速度均匀分布于0.25–1.5 m/s),其余回合为静态。
- 演示来自基于特权模拟器状态的脚本状态机,仅保留成功尝试。
- 单个物体(
can11),5个容器。 - 控制频率250 Hz,使用Panda机械臂,采用差分IK + 刚性关节位置控制器(
IK_KP=2500),实测末端执行器速度上限为1.11 m/s(与最快罐头相当,因此采用拦截而非追逐策略)。
数据结构
- 动作(
action)维度为(10,),包含:[dx, dy, dz, 3个欧拉角的正弦/余弦, 夹爪](共10维) - 状态(
observation.state)维度为(9,),包含:[x, y, z, 3个欧拉角的正弦/余弦](共9维) - 环境状态(
observation.environment_state)维度为(9,),包含特权物体位姿/速度,训练时应丢弃。 - 动作为增量式(delta):
绝对目标位姿 - 当前帧状态。评估时解码为目标 = 实时状态 + 预测增量,每步重新锚定。 - 传感器配置:3个RGB相机(
wrist_cam、side_cam、opst_cam),分辨率360×480,并配有各相机的v2e事件极性渲染。训练使用opst_cam与wrist_cam。 - 元数据:
meta/camera.jsonl将每个episode_index映射回源HDF5文件名。由于转换分片采用轮询方式,回合顺序并非源顺序,此文件是唯一的回溯途径。
前瞻偏移量的影响
偏移量决定了绝对目标所代表的含义:
futee0:标签为状态机在时间t命令的位姿(即控制器实际跟踪误差)。futee>0:标签为手臂在N帧后达到的位姿(即该窗口内的位移)。
在18个预留场景上的标签回放测试结果(任何基于这些标签训练的模型的理论上限):
| 偏移量 | 前瞻时间 | 命令步长(中位数) | 与状态机真实跟踪误差的比率 | 标签回放成功率 |
|---|---|---|---|---|
| 0 | 0 ms | 4.2 cm | 1.00× | 18/18 |
| 5 | 20 ms | ~1 cm | 0.2× | 0/18 |
| 10 | 40 ms | ~2 cm | 0.4× | 0/18 |
| 20 | 80 ms | 2.9 cm | 0.69× | 12/18 |
| 30 | 120 ms | — | — | 6/18 |
| 40 | 160 ms | — | — | 1/18 |
| 50 | 200 ms | 7.7 cm | 1.83× | 0/18 |
| 60 | 240 ms | — | — | 0/18 |
| 80 | 320 ms | 11.4 cm | 2.71× | 0/18 |
分析:
- 偏移过大会导致手臂需要在4 ms内闭合11 cm的差距,导致运行速度远超演示路径并产生过冲(静态场景轨迹显示,在回放步骤100时就到达专家的帧556,随后在20–32 cm处振荡)。
- 偏移过小则导致手臂爬行且时间不足。
- 只有
futee0能重现演示,因此推荐使用该变体。 - 直接回放源HDF5的绝对状态机动作在相同场景上的成功率为12/12,说明场景、物理和成功标准并非限制因素。
免训练可学习性检查:在物理情境几乎相同的不同回合帧中,标签相对其整体分布的离散度为:futee0时为0.001,futee20时升至0.178,320 ms时升至0.237——窗口越长,标签越依赖未来物理状态,而这些信息无法从观测中获得。
注意:以上数据是250 Hz控制下的特定结论。在25 fps下,200 ms的前瞻恰好产生4.0 cm的命令步长(对应4.1 cm跟踪误差的0.97×),因此25 fps配方使用--futee_offset 5并无问题。
复现方法
生成、转换、训练和评估脚本位于:https://github.com/mickeykang16/DynamicVLA/tree/mujoco



