longhorizon-subtask-decomposition
收藏资源简介:
该数据集用于长视野子任务分解任务。给定一个桌面操作场景的初始帧图像以及一条复合指令,模型需要生成完成该指令所需的一系列有序的单个拾取放置子任务。每个样本对应一个场景。数据集的帧图像和标签来自三个基于 SimplerEnv(WidowX/Bridge)、RobotArena(WidowX)和 RoboLab(Franka)构建的长视野组合基准套件。每个样本的帧图像是评估时编排器实际看到的重置观测图像,具有相同的相机视角、确定性布局和指令字符串。标签是真实标签,来源于每个套件的任务注册表,并经过严格校验,确保子任务序列与计划缓存中的真实计划一致。数据集包含102个场景,其中23个场景的任务需要按顺序执行(ordered),每个场景的目标数量(goals)从1到10不等。任务类型涵盖单一目标、分割目标、转移、接力、重访、排除、有序和重复等,使得任务长度和类型可以独立变化。字段包括:id(基准套件/场景)、benchmark(来源基准套件)、instruction(复合指令)、subtasks(有序子任务列表)、n_goals(目标数量,接力任务中可能不等于子任务数)、ordered(是否必须按顺序完成)、family(RoboLab任务族,其他套件为空)、distractors(干扰物对象)、image_main(重置帧图像)、image_wrist(腕部相机图像,仅RoboLab有)。数据集中61个样本带有帧图像,其余41个暂仅有标签。
This dataset is used for long-horizon subtask decomposition tasks. Given an initial frame image of a desktop manipulation scene and a composite instruction, the model needs to generate a series of ordered single pick-and-place subtasks required to complete the instruction. Each sample corresponds to a scene. The frame images and labels come from three long-horizon composite benchmark suites built on SimplerEnv (WidowX/Bridge), RobotArena (WidowX), and RoboLab (Franka). The frame image of each sample is the reset observation image actually seen by the orchestrator during evaluation, with the same camera viewpoint, deterministic layout, and instruction string. The labels are ground truth, sourced from the task registry of each suite, and strictly verified to ensure that the subtask sequence matches the real plan in the plan cache. The dataset contains 102 scenes, of which 23 scenes require tasks to be executed in order (ordered), and the number of goals per scene ranges from 1 to 10. Task types include single goal, split goal, transfer, relay, revisit, exclusion, ordered, and repetition, allowing task length and type to vary independently. Fields include: id (benchmark/scene), benchmark (source benchmark suite), instruction (composite instruction), subtasks (ordered list of subtasks), n_goals (number of goals, which may not equal the number of subtasks in relay tasks), ordered (whether it must be completed in order), family (RoboLab task family, empty for other suites), distractors (distractor objects), image_main (reset frame image), image_wrist (wrist camera image, only available for RoboLab). 61 samples in the dataset come with frame images, while the remaining 41 currently only have labels.
数据集概述:Long-horizon subtask decomposition
基本信息
- 任务类型:image-text-to-text(图像-文本到文本)
- 语言:英语(en)
- 标签:机器人学、视觉语言模型、任务分解、长程操作
- 数据规模:共 102 条样本(n<1K),全部为训练集
- 许可证:其他(other)
- 下载大小:约 20.7 MB
数据集内容
该数据集用于长程桌面操作场景的子任务分解任务。给定桌面操作场景的初始帧和复合指令,模型需生成完成该指令的有序单项拾取-放置子任务列表。每条数据对应一个场景。
数据来源包括三个基于 SimplerEnv(WidowX/Bridge)、RobotArena(WidowX)和 RoboLab(Franka)构建的长程组合套件。每行的帧为推理时编排器实际看到的重置观测(相同相机、确定性布局、相同指令字符串)。
数据标注
- 标签为真实值(ground truth),非模型输出。每个套件带有跟踪计划缓存,其
_meta.role为ground truth。 - 所有计划均源自该套件的任务注册表并经过校验。数据集的
subtasks列为这些计划的逐字记录。 - 模型自身的分解结果将根据这些标签进行评分,且不会与标签合并。
数据分布
| 基准(Benchmark) | 场景数 | 有序场景数 | 每场景目标数 |
|---|---|---|---|
| simplerenv_comp | 54 | 15 | 1–10 |
| robotarena_comp | 23 | 4 | 2–8 |
| robolab_comp | 25 | 4 | 1–10 |
| 总计 | 102 | 23 |
任务类型涵盖单一目的地、分离目的地、转移、中继、重访、排除、有序和重复等,长度和类型可独立变化。
数据字段
| 字段 | 含义 |
|---|---|
id |
<benchmark>/<scene> |
benchmark |
所属基准套件(simplerenv_comp、robotarena_comp、robolab_comp) |
instruction |
策略读取的复合指令 |
subtasks |
真实值的有序子任务列表 |
n_goals |
场景评分的目标数。除中继任务(需通过中间容器进行两次放置)外,等于 subtasks 长度 |
ordered |
目标是否必须按列出顺序完成。对于后段在前段之前不可能完成的任务链(中继、重访)以及指令中明确顺序的任务为 True |
family |
RoboLab 任务族(single_dest、split、relay、transfer 等);其他套件为空 |
distractors |
存在但不属于任何目标的对象 |
image_main |
重置帧,即编排器看到的画面 |
image_wrist |
腕部相机(仅 RoboLab 套件) |
帧覆盖情况
- 102 行中有 61 行带有帧,其余 41 行目前仅包含标签:其中 11 个 SimplerEnv 场景和 5 个 RobotArena 场景为新增场景,而全部 25 个 RoboLab 场景在之前版本后基于测量资产重建了布局,导致旧帧失效。
- 帧仅在场景布局输入被证明未改变时才会被复用(包括对象、容器、尺度、生成覆盖和抖动),因此每行要么显示正确的帧,要么不显示帧,绝不会有过期帧。
- 未覆盖的帧将由各套件的
export_scenes重新渲染,并在后续修订版本中补充。





