EliGen Dataset
收藏资源简介:
EliGen数据集由浙江大学和阿里巴巴集团联合创建,旨在支持实体级控制的图像生成任务。该数据集包含50万条高质量的训练样本,每个样本包括图像、全局提示、局部提示和实体掩码。数据集的生成过程使用了Flux生成图像,并通过Qwen2-VL视觉语言模型进行全局提示和实体信息的标注。该数据集的应用领域主要集中在图像生成和图像修复任务,旨在解决现有文本到图像生成模型在细粒度实体控制上的不足,提供更精确的实体位置和语义控制能力。
The EliGen dataset, jointly developed by Zhejiang University and Alibaba Group, is designed to support entity-level controlled image generation tasks. It contains 500,000 high-quality training samples, each comprising an image, a global prompt, a local prompt, and an entity mask. In the dataset generation process, Flux is utilized to generate images, and the Qwen2-VL vision-language model is employed to annotate global prompts and entity information. Its application domains mainly focus on image generation and image inpainting tasks, aiming to address the shortcomings of existing text-to-image generation models in fine-grained entity control, and provide more precise control over entity positions and semantics.




