qwen-deepfashion
收藏资源简介:
Qwen DeepFashion(真实+合成全身服装)数据集包含160,015张完全由AI生成的全身时尚图像,使用Qwen-Image模型和Qwen-Image-Lightning 4步LoRA生成。服装描述来源于两个来源:真实的DeepFashion描述集和合成的服装生成器。通过共享的提示增强策略(fashion-v1),将这些描述渲染为从头到脚的服装照片,同时多样化穿着者、背景、姿势和构图,并保留每个源描述的性别适宜服装。该数据集旨在作为全身时尚生成的广泛、平衡的预训练基础,补充其肖像导向的姊妹数据集。数据集包含图像、提示文本和多种元数据字段,如人口统计意图(种族、性别、年龄区间等)、图像尺寸和质量标记。数据组成显示女性主导(70.5%),种族分布经过平衡采样,背景以真实世界设置为主(82%),图像分辨率包括1024×1024、832×1216和1216×832三种比例。该数据集适用于生成图像模型的研究与开发,特别是全身/时尚生成任务,包括预训练和微调数据、数据增强、服装和姿势条件生成,以及研究合成时尚数据中的人口统计平衡。重要注意事项包括:所有图像均为合成生成,不描绘真实人物;人口统计列仅表示生成意图,非真实属性;使用前需进行年龄验证过滤;数据集设计上女性占主导地位。
The Qwen DeepFashion (real + synthetic full-body clothing) dataset contains 160,015 fully AI-generated full-body fashion images, generated using Qwen-Image and Qwen-Image-Lightning 4-step LoRA. Clothing descriptions are derived from two sources: the real DeepFashion description dataset and a synthetic clothing generator. Via a shared prompt augmentation strategy (fashion-v1), these descriptions are rendered into head-to-toe clothing photos, with diversified wearers, backgrounds, poses and compositions, while retaining gender-appropriate clothing for each source description. This dataset aims to serve as a comprehensive and balanced pre-training foundation for full-body fashion generation, complementing its portrait-oriented sibling dataset. The dataset includes images, prompt texts, and multiple metadata fields such as demographic intent (race, gender, age range, etc.), image dimensions and quality markers. The dataset composition shows a female-majority distribution (70.5%), with a balanced sampled racial distribution; backgrounds are predominantly real-world settings (82%), and the image resolutions include three aspect ratios: 1024×1024, 832×1216 and 1216×832. This dataset is applicable to research and development of generative image models, particularly full-body/fashion generation tasks, including pre-training and fine-tuning data, data augmentation, clothing and pose-conditioned generation, and research on demographic balance in synthetic fashion data. Important notes include: all images are synthetically generated and do not depict real individuals; demographic columns only represent generation intent rather than real attributes; age verification filtering is required before use; the dataset is intentionally female-majority in its composition.
数据集概述
Qwen DeepFashion (Real + Synthetic Full-Body Outfits) 是一个包含 160,015 张 全身体验时尚图像的合成数据集,由 Qwen-Image 和 Qwen-Image-Lightning 模型生成。数据集旨在为全身体验时尚生成任务提供广泛、平衡的预训练数据。
数据来源与构成
- 图像: 全部为 AI 生成的合成图像,分辨率为 1024×1024 (35.3%)、832×1216 (33.4%)、1216×832 (31.2%)。
- 提示词来源:
- 真实 DeepFashion 描述 (约 12,015 条): 源自
AbstractPhil/diffusion-pretrain-set-ft1数据集的 DeepFashion 文本描述(仅文本,不包含原始图像)。 - 合成服装描述 (约 148,000 条): 由性别区分的组合服装生成器创建,用于平衡 DeepFashion 中某些服装类别(如连衣裙、正装)的不足。合成部分女性占多数(约 7:3),同时包含男性平行数据。
- 真实 DeepFashion 描述 (约 12,015 条): 源自
生成过程
- 模型: Qwen-Image (~20B) + Qwen-Image-Lightning (4 步 LoRA)。
- 推理参数: 4 步推理,
true_cfg_scale=1.0,bf16 精度。 - 提示增强策略 (policy_version = fashion-v1):
- 保留源服装描述及其性别信息。
- 丰富场景:约 18% 为工作室背景,约 82% 为真实世界场景。
- 确保每张图像为全身、从头到脚的构图。
- 注入多样化元素:姿势 (约 75%)、拍摄视角 (约 55%)、相机角度 (约 50%)、光照 (约 60%)。
- 人口统计学平衡: 当源描述未指定种族时,从近似均匀分布中采样标签,并标注长尾标签 (
is_tail)。 - 年龄限制: 强制主题年龄为 25-35 岁成年,并去未成年化处理。
- 质量分级: 约 90% 为高质量彩色摄影,约 10% 为故意低质量的“业余”风格 (
is_amateur)。 - 去重机制: 使用基于 MinHash/LSH 的
DiversityGuard机制,阈值 0.82,防止提示词过度相似。
数据集结构
数据集中包含以下列:
| 列名 | 类型 | 说明 |
|---|---|---|
id |
string | 稳定唯一键;deepfashion_ 前缀代表真实描述,fsyn_ 前缀代表合成数据 |
image |
Image | 生成的 PNG 格式图像 |
image_width, image_height |
int32 | 图像像素尺寸 |
prompt |
string | 增强后用于生成的提示词 |
source_prompt |
string | 增强前的原始服装描述 |
race |
string | 意图的种族/民族标签(生成意图,非验证属性) |
race_injected |
bool | 种族标签是否为采样获得 (true) 或源自源描述 (false) |
is_tail |
bool | 种族标签是否属于长尾类别 |
gender |
string | 意图的性别:woman、man、person |
age_band |
string | 始终为 25-35 |
hair |
string | 意图的发型属性(约 70% 行注入,其余为空) |
eye, expression, makeup, jewelry |
string | 设计上为空(仅限肖像字段,全身体验策略中未使用) |
is_amateur |
bool | 是否为“业余”风格 (约 10%) |
seed |
int64 | 每行唯一的确定性种子 |
width_ratio |
string | 如 832x1216 |
policy_version |
string | 增强策略版本 (fashion-v1) |
数据组成统计
- 来源: 合成女性 101,000 (63.1%),合成男性 47,000 (29.4%),真实 DeepFashion 12,015 (7.5%)。
- 意图性别: 女性 112,785 (70.5%),男性 47,018 (29.4%),无标签 212 (0.1%)。
- 意图种族: 白种人 (14.9%) 最多,其余主要类别各约 8.0%,长尾类别各约 2.0%。99.9% 的标签为采样生成。
- 背景: 现实场景 82.0%,工作室 17.9%。
- 构图: 所有图像均为全身/从头到脚构图。
- 质量: 业余级约 10.0%。
- 合成服装类别 (148,000 行):
- 女装:连衣裙 (20,259)、外套 (17,108)、针织 (14,032)、正装 (13,956) 等。
- 男装:上衣 (8,805)、正装 (7,220)、下装 (7,022) 等。
- 去重率: 精确重复提示词率 0.005%,归一化重复率 0.006%。
- 图像饱和度: 无灰度或接近灰度的图像。
预期用途
主要用于全身体验时尚生成类生成式图像模型的研究与开发,包括预训练、微调、数据增强、服装和姿态条件生成,以及合成时尚数据中人口统计学平衡的研究。
限制与注意事项
- 年龄 / 未成年人: 提示词虽已限制为成人,但未进行自动年龄验证。模型可能生成外貌偏年轻的图像。用户在使用前必须执行严格的年龄验证并移除任何可能描绘未成年人的样本。
- 标签 ≠ 真实: 人口统计学和属性列仅为生成意图,并非对生成图像的验证,不可用于训练或评估真实人物的分类器。
- 背景 / 服装统计: 背景统计接近准确,服装类别统计仅对合成行可靠,真实 DeepFashion 描述为自由文本,不在此统计范围内。
- 图像偏差: 数据集女性占主导 (约 7:3),合成行中存在约 10% 的刻意低质量图像,且可能含有合成痕迹。
- 许可: 生成组件基于 Apache-2.0 许可。数据集本身标记为
apache-2.0,但用户应根据自身用途确认最终许可,尤其是商业用途时需注意 DeepFashion 原始条款。




