GML-FMGroup/CanvasCraftRL
收藏资源简介:
CanvasCraftRL是CanvasCraft数据集的强化学习任务规范子集,最初在“CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration”中引入。它专为训练和评估多模态代理而设计,这些代理通过在多轮中协调多个视觉工具来解决复杂的视觉创建和编辑请求。与监督轨迹数据集不同,CanvasCraftRL不提供固定的推理痕迹、工具顺序、参数序列、中间观察或最终输出。每个示例提供用户任务、可选输入图像和预期工具集。这种弱监督允许强化学习方法(如GRPO)探索替代工具使用策略,同时仍从预期工具中获得过程级指导。数据集包含9,110个训练行和250个测试行,总计9,360行,并附带8,745个归一化PNG图像。
CanvasCraftRL is the reinforcement-learning task-specification subset of the CanvasCraft dataset introduced with CanvasAgent: Enabling Complex Image Creation and Editing via Visual Tool Orchestration. It is designed for training and evaluating multimodal agents that solve complex visual creation and editing requests by orchestrating multiple visual tools over several turns. Unlike supervised trajectory datasets, CanvasCraftRL does not prescribe a fixed reasoning trace, tool order, parameter sequence, intermediate observation, or final output. Each example provides a user task, optional input images, and an expected tool set. This weak supervision allows RL methods such as GRPO to explore alternative tool-use strategies while still receiving process-level guidance from the expected tools. The dataset contains 9,110 training rows and 250 test rows, totaling 9,360 rows, accompanied by 8,745 normalized PNG images.




