MeiGen-AI/GenEvolve-Data-Bench
收藏资源简介:
GenEvolve数据集是一个用于图像生成和视觉问答的多模态基准数据集,包含三个主要配置:sft(监督微调)提供9,000个轨迹和50,291张参考图像,用于监督冷启动轨迹;rl(强化学习)提供3,175个提示和对应GT图像,用于自进化或RL训练;bench(基准测试)提供594个提示和GT图像,用于保留评估。数据集支持文本到图像生成、视觉轨迹分析和代理学习任务,适用于多模态AI研究,如工具编排的视觉经验蒸馏。数据以JSONL和Parquet格式提供,图像来自公共网络资源或合成轨迹,遵循Apache-2.0许可证。
The GenEvolve dataset is a multimodal benchmark dataset for image generation and visual question answering. It comprises three primary configurations: the SFT (Supervised Fine-Tuning) configuration supplies 9,000 trajectories and 50,291 reference images for supervised cold-start trajectory training; the RL (Reinforcement Learning) configuration provides 3,175 prompts and their corresponding ground-truth (GT) images for self-evolution or RL training; and the Bench configuration offers 594 prompts and GT images for held-out evaluation. This dataset supports text-to-image generation, visual trajectory analysis, and agent learning tasks, and is applicable to multimodal AI research including visual experience distillation via tool orchestration. The data is provided in JSONL and Parquet formats, with images sourced from public web resources or synthetic trajectories, and it is released under the Apache-2.0 license.



