SeFi-Image/SeFi-Image-DiT-Finetune-Demo
收藏资源简介:
该数据集包含56个训练和8个验证图像-文本对,用于固定1024 DiT(扩散变换器)微调的烟雾测试。数据来源于ma-xu/fine-t2i数据集的synthetic_enhanced_prompt_random_resolution子集,仅包含图像生成器为Z-Image-Turbo的行,并排除了FLUX.2-dev行、现有SemVAE演示源ID以及在提示/图像质量审查中被拒绝的行。数据集通过稳定的源ID进行筛选,并确定性地以4:1的权重选择增强提示和原始提示。图像为至少1024像素的正方形源JPEG文件,在训练加载器中会确定性地调整为1024x1024分辨率。每个数据行包含嵌入的图像、提示、增强提示、来源信息、尺寸、图像SHA-256哈希、生成器元数据以及ai_generated=true标记。该数据集是一个小型合成工程测试数据集,不具代表性训练语料库或模型质量基准功能,生成场景可能包含不完美的解剖结构或渲染文本。
This dataset contains 56 training and 8 validation image-text pairs for fixed-1024 DiT fine-tuning smoke tests. It is derived from the ma-xu/fine-t2i datasets synthetic_enhanced_prompt_random_resolution subset, including only rows where the image generator is Z-Image-Turbo, and excludes FLUX.2-dev rows, existing SemVAE demo source IDs, and rows rejected during prompt/image quality review via stable source ID. The dataset selects enhanced_prompt and prompt with deterministic 4:1 weights, and images are square source JPEGs at least 1024px, resized deterministically to 1024x1024 in the training loader. Each row contains an embedded image, prompt, enhanced_prompt, source provenance, dimensions, an image SHA-256, generator metadata, and an ai_generated=true marker. It is a tiny synthetic engineering fixture, not a representative training corpus or model-quality benchmark, and generated scenes may contain imperfect anatomy or rendered text.




