遇见数据集

novastar112/toy_maze_2d_hard_allstep_thinking_future_rollout_cot_500k

收藏
Hugging Face2026-05-09 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是从本地VisGym ToyMaze2D的maze_2d/hard环境中生成的。它包含训练集(500,000行gzip压缩的JSONL格式数据)和测试集(100行gzip压缩的JSONL格式数据)。每行代表一个完整的轨迹对话:每个用户回合存储一个提示和三个JPEG图像项(image_prev、image和image_next),其中每一步都验证image_prev == image,最终步骤有image_next == image。在轨迹中,两个非停止移动步骤包含四个一步未来滚动图像(右、上、左、下),且loss=true;其他非停止步骤使用固定的简单思考模板。最终停止步骤使用固定的停止思维链(COT)并输出(stop, stop)。提示步骤预算为40个总步骤,用户提示中显示剩余步骤为max_steps - step_index - 1。

This dataset is generated from the local VisGym ToyMaze2D maze_2d/hard environment. It includes a training set (500,000 gzip-compressed JSONL rows) and a test set (100 gzip-compressed JSONL rows). Each row is a full trajectory conversation: every user turn stores a prompt and three JPEG image items (image_prev, image, and image_next), with image_prev == image validated for every step, and the final step having image_next == image. Two non-stop move steps per trajectory contain four one-step future rollout images (right, up, left, down) with loss=true; other non-stop steps use a fixed trivial thinking template. The final stop step uses a fixed stop Chain of Thought (COT) and outputs (stop, stop). The prompt step budget is 40 total steps for hard, and the user prompt indicates remaining steps as max_steps - step_index - 1.

提供机构:
novastar112
二维码
社区交流群
二维码
科研交流群
商业服务