GML-FMGroup/CanvasCraftSFT
收藏资源简介:
CanvasCraftSFT是CanvasCraft数据集的监督微调子集,源自CanvasAgent项目,专注于复杂图像创建和编辑任务的可执行多模态工具使用轨迹。该数据集旨在训练代理通过用户请求进行推理、调用带有结构化参数的视觉工具、观察中间视觉结果,并决定图像转换何时完成。数据集包含138,990个轨迹示例,涉及166,563个嵌入图像和2,000个额外SR图像,覆盖多种工具如ImageEdit、ImageGeneration、OCR等。数据格式包括JSON轨迹文件(包含任务提示、图像列表和聊天式消息)和Parquet图像分片(包含图像字节和路径)。数据集主要用于多模态工具使用代理的监督微调研究,支持图像创建和编辑工作流程、视觉代理的推理-动作-观察训练,以及为CanvasCraftRL的强化学习阶段提供基础。局限性包括路径前缀需标准化、图像存储方式不同以及工具环境固定。
CanvasCraftSFT is the supervised fine-tuning subset of the CanvasCraft dataset introduced with CanvasAgent, focusing on executable multimodal tool-use trajectories for complex image creation and editing tasks. It teaches agents to reason over user requests, call visual tools with structured arguments, observe intermediate visual results, and decide when image transformations are complete. The dataset contains 138,990 trajectory examples, with 166,563 embedded images and 2,000 additional SR images, covering tools such as ImageEdit, ImageGeneration, and OCR. The data format includes JSON trajectory files (with task prompts, image lists, and chat-style messages) and Parquet image shards (with image bytes and paths). It is intended for research on supervised fine-tuning of multimodal tool-use agents, image creation and editing workflows, reasoning-action-observation training for visual agents, and bootstrapping before RL optimization on CanvasCraftRL. Limitations include the need to normalize legacy path prefixes, mixed image storage, and a fixed tool environment.




