UniWorld-V1
收藏资源简介:
UniWorld数据集是一个用于图像理解和生成任务的统一数据集,由北京大学深圳研究生院和鹏城实验室等机构创建。该数据集包含约2.7M个样本,包括图像感知和操作任务的数据,例如检测、分割、深度预测、添加、调整、提取等。数据集的创建过程包括使用高质量的开源数据、自生成数据和过滤后的开源数据,并使用自适应编辑区域加权策略来处理图像编辑任务。UniWorld数据集旨在解决图像理解和生成任务中的挑战,并支持多模态领域的研究和开发。
The UniWorld Dataset is a unified dataset for image understanding and generation tasks, developed by institutions including Peking University Shenzhen Graduate School and Peng Cheng Laboratory. This dataset contains approximately 2.7 million samples, covering data for image perception and manipulation tasks such as detection, segmentation, depth prediction, image addition, adjustment, and extraction. The dataset's development process incorporates high-quality open-source data, self-generated data, and filtered open-source data, and adopts an adaptive editing region weighting strategy to handle image editing tasks. The UniWorld Dataset aims to address the challenges in image understanding and generation tasks, and supports research and development in the multimodal domain.
数据集概述
基本信息
- 许可证: MIT
- 相关论文: UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
- 更多详情: UniWorld-V1
数据来源
- Geneval-style数据集: 来源于BLIP3o-60k,其中一半数据添加了文本到图像的指令。[108 GB存储空间]
数据分类及详情
文本到图像生成
- BLIP3o-60k: 添加了文本到图像的指令。[108 GB存储空间]
- OSP1024-286k: 来源于Open-Sora Plan内部数据,使用Qwen2-VL-72B生成标题。图像宽高比在3:4到4:3之间,美观度评分≥6,短边≥1024像素。[326 GB存储空间]
图像编辑
- imgedit-724k: 使用GPT-4o过滤,保留约一半数据。[2.8T存储空间]
- OmniEdit-368k: 过滤掉编辑区域小于1/100的样本,图像短边≥1024像素。[204 GB存储空间]
- SEED-Data-Edit-Part1-Openimages-65k: 过滤掉编辑区域小于1/100的样本,图像短边≥1024像素。[10 GB存储空间]
- SEED-Data-Edit-Part2-3-12k: 过滤掉编辑区域小于1/100的样本,图像短边≥1024像素。[10 GB存储空间]
- PromptfixData-18k: 用于图像修复和部分编辑数据,过滤掉编辑区域小于1/100的样本,图像短边≥1024像素。[9 GB存储空间]
- StyleBooth-11k: 用于风格转换数据,图像短边≥1024像素。[4 GB存储空间]
- Ghibli-36k: 用于风格转换数据,图像短边≥1024像素。警告:此数据未经过质量过滤。[170 GB存储空间]
提取与试穿
- viton_hd-23k: 从源数据转换为产品提取的指令数据集。[1 GB存储空间]
- deepfashion-27k: 从源数据转换为产品提取的指令数据集。[1 GB存储空间]
- shop_product-23k: 来源于Open-Sora Plan内部数据,专注于产品提取和虚拟试穿,图像短边≥1024像素。[12 GB存储空间]
图像感知
- coco2017_caption_canny-236k: 图像到canny边缘检测及反向操作。[25 GB存储空间]
- coco2017_caption_depth-236k: 图像到深度图及反向操作。[8 GB存储空间]
- coco2017_caption_hed-236k: 图像到HED边缘检测及反向操作。[13 GB存储空间]
- coco2017_caption_mlsd-236k: 图像到MLSD边缘检测及反向操作。[存储空间未指定]
- coco2017_caption_normal-236k: 图像到法线图及反向操作。[10 GB存储空间]
- coco2017_caption_openpose-62k: 图像到姿态估计及反向操作。[2 GB存储空间]
- coco2017_caption_sketch-236k: 图像到草图及反向操作。[15 GB存储空间]
- unsplash_canny-20k: 图像到canny边缘检测及反向操作。[2 GB存储空间]
- open_pose-40k: 图像到姿态估计及反向操作。[4 GB存储空间]
- mscoco-controlnet-canny-less-colors-236k: 图像到canny边缘检测及反向操作。[13 GB存储空间]
- coco2017_seg_box-448k: 图像到检测和分割(掩码),过滤掉区域小于1/100的实例。[39 GB存储空间]
- viton_hd-11k: 图像到姿态估计。[1 GB存储空间]
- deepfashion-13k: 图像到姿态估计。[1 GB存储空间]

- 1UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation北京大学深圳研究生院 · 2025年



