Uni-Edit-148k
收藏资源简介:
Uni-Edit-148k是由香港中文大学多媒体实验室等机构构建的智能图像编辑数据集,旨在通过单一任务统一提升多模态模型的理解、生成与编辑能力。该数据集包含14.8万条高质量样本,每条数据由复杂的推理密集型编辑指令与对应编辑后的图像组成,其指令源自LLaVA-OneVision-1.5的视觉问答数据,并利用Nano-Pro生成目标图像,再经GPT-4o严格筛选确保逻辑一致性与视觉美感。该数据集主要应用于多模态模型的统一调优,通过将视觉理解任务转化为嵌入式逻辑的编辑指令,有效解决了传统多任务训练中理解与生成能力冲突的难题,为模型实现全面协同增强提供了创新解决方案。
Uni-Edit-148k is an intelligent image editing dataset developed by the Multimedia Laboratory of The Chinese University of Hong Kong and other institutions, aiming to uniformly enhance the understanding, generation, and editing capabilities of multimodal models via a single task. This dataset contains 148,000 high-quality samples, each consisting of complex reasoning-intensive editing instructions and their corresponding edited images. The editing instructions are derived from the visual question answering data of LLaVA-OneVision-1.5, and target images are generated using Nano-Pro, followed by strict screening via GPT-4o to ensure logical consistency and visual aesthetics. This dataset is primarily applied to the unified fine-tuning of multimodal models. By transforming visual understanding tasks into editing instructions with embedded logic, it effectively resolves the conflict between understanding and generation capabilities in traditional multi-task training, providing an innovative solution for comprehensive collaborative enhancement of models.




