遇见数据集

MultiEdit

收藏
魔搭社区2026-09-06 更新2026-09-06 收录
官方服务:

资源简介:

# 🧩 MultiEdit: Advancing Instruction-based Image Editing on Diverse and Challenging Tasks [📃 Arxiv](https://arxiv.org/abs/2509.14638) --- ## 🚀 Dataset Overview Based on our MLLM-driven data construction pipeline using GPT-4o and GPT-Image-1, we introduce **MultiEdit**, a comprehensive large-scale instruction-based image editing dataset comprising over 107K samples targeting **6 challenging image editing tasks** covering **56 subcategory editing types** (**18 non-style-transfer** and **38 style transfer**). We also release MultiEdit-Test, a carefully curated benchmark of 1.1K samples to assess complex editing capabilities. The involved 6 image editing tasks are as follows: - 🖼️ **Object Reference Editing**: Modifies specific attributes (color, shape, scale, and position) of referenced objects. - 👤 **Person Reference Editing**: Targets referenced individuals within images, altering their pose, clothing, hairstyle, skin color, and figure. - ✍️ **Text Editing**: Focuses on textual elements within movie posters, covering modifications in font style, expression, display medium, and font color. - 📱 **GUI Editing**: Modifies icon attributes and the display medium of GUI elements, using images of diverse digital interfaces (e.g., iOS, Android, and websites). - 👁️ **View Editing**: Generates alternative views of subjects within images, encompassing edits for persons, landmarks, and general objects. - 🎨 **Style Transfer**: Reimagines images with 38 distinct artistic styles, from classical art to modern digital aesthetics. ## 📊 Dataset Statistics The following table provides a detailed statistical breakdown of the MultiEdit dataset by task, including the number of edit types and the distribution of samples between the training and test sets. | Task | # of Edit Types | Train Samples | Test Samples | Total Samples | | -------------------------- |:---------------:| ------------- | ------------ | ------------- | | Object Reference Editing | 4 | 9,851 | 200 | 10,051 | | Person Reference Editing | 5 | 6,891 | 250 | 7,141 | | Text Editing | 4 | 3,860 | 200 | 4,060 | | GUI Editing | 2 | 2,780 | 100 | 2,880 | | View Editing | 3 | 28,055 | 150 | 28,205 | | Style Transfer | 38 | 55,097 | 200 | 56,297 | | **Total** | **56** | **106,534** | **1,100** | **107,634** | ## 🏗️ Data Structure The organization of the MultiEdit-Train and MultiEdit-Test sets is defined by their respective `metadata.json` files. The unified structure of these JSONL files is as follows: ```text [ { "original_images": "XXX", // path to source image "generated_images": "XXX", // path to edited image "edit_prompt": "XXXXX", // the edit instruction "meta_prompt_index": X, // (Optional) index of edit type, corresponding to the order in Table 1 of our paper. "source": "XX", // the dataset source of the original image (e.g., 'GUI_World') "id": xxx, // a unique ID to index this data triplet } ] ``` ## 🤝 Acknowledgements We would like to thank the following research works and projects: - [GPT-4o](https://openai.com/index/gpt-4o-image-generation-system-card-addendum/) - [GPT-Image-1](https://openai.com/index/image-generation-api/) - [SD3](https://huggingface.co/stabilityai/stable-diffusion-3-medium) - [UltraEdit](https://ultra-editing.github.io/) ## 🧾 License [![License](https://img.shields.io/badge/license-Apache_2.0-green.svg)](#license) This project is licensed under the Apache-2.0 License. ## 📢 Citation If you find our work useful for your research, please consider citing our paper: ```bibtex @article{li2025multiedit, title={MultiEdit: Advancing Instruction-based Image Editing on Diverse and Challenging Tasks}, author={Mingsong Li and Lin Liu and Hongjun Wang and Haoxing Chen and Xijun Gu and Shizhan Liu and Dong Gong and Junbo Zhao and Zhenzhong Lan and Jianguo Li}, journal={arXiv preprint arXiv:2509.14638}, year={2025}, } ```

提供机构:
maas
创建时间:
2026-08-31
二维码
社区交流群
二维码
科研交流群
商业服务