unipic_seedream_4images
收藏资源简介:
UniPic-Nano-4Images 是一个高质量的多图像合成数据集,包含 48,805 个样本,专为训练先进的图像融合和合成模型而设计。每个样本由 4 张输入图像和 1 张输出图像组成,根据自然语言指令将四张输入图像中的元素无缝组合。该数据集是 UniPic 系列的一部分,已用于 UniPic3 中训练具有多元素融合能力的多图像合成模型。 数据集特点包括: - 4 图像输入:每个样本使用 4 张输入图像进行多元素合成 - 多元素融合:以复杂方式结合人物与 3 个对象/元素 - 多样化合成模式:涵盖多种同时动作的合成场景 - 高质量:48,805 个精心筛选的样本,附带详细自然语言指令 - 生产就绪:已用于实际多图像合成应用 - 简单格式:清晰的 JSON 格式,结构简单明了 数据集统计包括动作分布(如持有、穿着、站立等)、动作组合分布(如持有+站立+穿着)和元素类型分布(如对象、穿戴物、家具等)。数据集采用一致的 4 元素合成模式:[来自 Image1 的主体] + [来自 Image2-4 的元素] → [融合输出]。 该数据集适用于训练和评估: - 高级多图像合成模型 - 复杂场景理解模型 - 遵循指令的视觉模型 - 多元素融合技术
UniPic-Nano-4Images is a high-quality multi-image synthesis dataset containing 48,805 samples, specifically designed for training advanced image fusion and synthesis models. Each sample consists of 4 input images and 1 output image, which seamlessly integrates elements from the four input images based on natural language instructions. This dataset is part of the UniPic series, and has been utilized to train multi-image synthesis models with multi-element fusion capabilities in UniPic3. Dataset features are as follows: - 4-image input: Each sample employs 4 input images for multi-element synthesis - Multi-element fusion: Combines human subjects and 3 objects/elements in complex scenarios - Diverse synthesis modes: Covers a wide range of synthesis scenarios involving simultaneous multiple actions - High-quality: 48,805 carefully screened samples accompanied by detailed natural language instructions - Production-ready: Has been deployed in practical multi-image synthesis applications - Simple format: Clear and straightforward JSON format with a concise and clear structure Dataset statistics include action distributions (e.g., holding, wearing, standing, etc.), action combination distributions (e.g., holding + standing + wearing), and element type distributions (e.g., objects, wearables, furniture, etc.). The dataset adopts a consistent 4-element synthesis paradigm: [Subject from Image 1] + [Elements from Images 2-4] → [Fused Output]. This dataset is applicable for training and evaluating: - Advanced multi-image synthesis models - Complex scene understanding models - Instruction-following visual models - Multi-element fusion technologies




