AbstractPhil/diffusion-pretrain-set-ft1
收藏资源简介:
diffusion-pretrain-set-ft1 是一个多源图像-文本预训练数据集,从十个上游来源(big_liminal、deepfashion、ffhq、flux_assorted_bulk、flux_assorted_bulk_2、imagenet_synthetic、imdb、mannequins_v7、mannequins_v10、synth_chars)通过统一的数据处理流程整合而成。该数据集专为扩散模型(如Stable Diffusion、ControlNet)的完整预训练或微调设计,旨在创建更强大的基线预训练数据,并为训练下一代视觉语言模型(VLM)提供图像合成基础。数据集包含结构化标注(JSON格式),提供两个并行标注列:caption_vlm_json(基于Qwen3.5-0.8B VLM从图像直接生成)和caption_animetimm_json(基于animetimm标签生成)。数据集经过多层过滤(包括年龄分类器和正则表达式),以移除不当内容,确保数据质量。数据规模在10万到100万之间,包含图像、条件图像、掩码及元数据,适用于多任务学习如文本到图像生成、控制网络训练等。
diffusion-pretrain-set-ft1 is a multi-source image-caption pretraining dataset assembled from ten upstream sources via a uniform ingest pipeline. Designed for a full pretrain or finetune pipeline for diffusion models like Stable Diffusion and ControlNet, it aims to create a more powerful baseline pretraining dataset and a foundation for synthesizing images to train the next generation of VLM models. The dataset features structured captions in JSON format, with two parallel caption columns: caption_vlm_json (generated directly from images using Qwen3.5-0.8B VLM) and caption_animetimm_json (generated from animetimm tags). It includes rigorous filtering (age classifier and regex) to remove inappropriate content. The size ranges between 100K and 1M samples, containing images, conditioning images, masks, and metadata, suitable for multi-task learning such as text-to-image generation and control network training.




