InternVL-U/ScaleEdit-12M
收藏资源简介:
--- license: mit task_categories: - image-to-image language: - en tags: - image-editing - instruction-based-editing - multimodal - computer-vision - scaleedit - internvl size_categories: - 10M<n<100M --- # ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework <div> [](https://arxiv.org/abs/2603.20644) [](https://github.com/gzchen4ai/ScaleEdit-12M) [](https://huggingface.co/datasets/InternVL-U/ScaleEdit-12M) </div> ## 📌 Overview **The largest open-source instruction-based image editing dataset to date.** ScaleEdit-12M contains **12.4 million** rigorously verified instruction–image pairs spanning **23 task families** across diverse real and synthetic visual domains. It was constructed using **ScaleEditor**, a fully open-source hierarchical multi-agent framework that eliminates the need for costly proprietary APIs.  ## 🔥 News - **[2026/04/03]** 🚀ScaleEdit-12M is released on [[Huggingface]](https://huggingface.co/datasets/InternVL-U/ScaleEdit-12M). - **[2026/03/24]** 🔥ScaleEdit-12M paper is released on [[arXiv]](https://arxiv.org/abs/2603.20644). - **[2026/03/06]** 🔥InternVL-U **technical report** released. Check it out on [[arXiv]](https://arxiv.org/abs/2603.09877). ## ✅ TODO - [x] Release ScaleEdit-12M dataset - [ ] Release ScaleEdit-1M subset - [ ] Release ScaleEditor framework ## 📊 Dataset Structure ### Repository Layout The dataset is organized into **23 task-specific subdirectories**, each containing multiple sharded Parquet files. The directory naming follows the pattern `{category_id}_{task_name}`: ``` ScaleEdit-12M/ ├── README.md ├── 1.1_style_transfer/ # Global editing tasks │ ├── style_transfer_0000.parquet # (~31.7 GB per shard) │ ├── style_transfer_0001.parquet │ ├── ... │ └── style_transfer_0015.parquet ├── 1.2_tone_adjustment/ │ └── tone_adjustment_XXXX.parquet ├── 1.3_viewpoint_transformation/ ├── 1.4_background_replacement/ ├── 2.1_object_addition/ # Object editing tasks ├── 2.2_object_removal/ ├── 2.3_object_replacement/ ├── 2.4_action_editing/ ├── 2.5_part_extraction/ ├── 3.1_color_change/ # Attribute editing tasks ├── 3.2_material_change/ ├── 3.3_visual_beautification/ ├── 3.4_count_change/ ├── 3.5_size_change/ ├── 4.1_movie_poster_text_editing/ # Text editing tasks ├── 4.2_gui_interface_text_editing/ ├── 4.3_object_surface_text_editing/ ├── 4.4_building_surface_text_editing/ ├── 5.1_perceptual_reasoning/ # Knowledge-infused tasks ├── 5.2_symbolic_reasoning/ ├── 5.3_social_reasoning/ ├── 5.4_scientific_reasoning/ └── 6.1_compositional_editing/ # Compositional tasks ``` Each task folder contains **multiple Parquet shards** (typically ~31–32 GB each) named `{task_name}_{shard_index:04d}.parquet`. The number of shards varies by task depending on the volume of data in that category. ### Parquet Schema Each Parquet file contains the following columns: | Column | Type | Description | |---|---|---| | `id` | `int64` | Unique identifier for the sample | | `edit_task` | `string` | Task category name (e.g., `"style_transfer"`, `"object_addition"`) | | `edit_instruction` | `string` | Natural-language editing instruction | | `source_image` | `binary` | Raw bytes of the source image (pre-edit) | | `edited_image` | `binary` | Raw bytes of the edited image (post-edit) | | `source_image_width` | `int64` | Width of the source image in pixels | | `source_image_height` | `int64` | Height of the source image in pixels | | `edited_image_width` | `int64` | Width of the edited image in pixels | | `edited_image_height` | `int64` | Height of the edited image in pixels | | `instruction_following_score` | `int64` | Quality score: how well the edit follows the instruction (1–3) | | `editing_consistency_score` | `int64` | Quality score: consistency between source and edited images (1–3) | | `generation_quality_score` | `int64` | Quality score: overall visual quality of the edited image (1–3) | ### Example Row ```json { "id": 0, "edit_task": "object_addition", "edit_instruction": "Add a red and white striped safety barrier at the edge of the platform on the right side of the image.", "source_image": <binary bytes>, "edited_image": <binary bytes>, "source_image_width": 2000, "source_image_height": 1500, "edited_image_width": 2000, "edited_image_height": 1500, "instruction_following_score": 3, "editing_consistency_score": 3, "generation_quality_score": 3 } ``` The `source_image` and `edited_image` columns store images as raw binary bytes. They can be decoded into PIL images: ```python from PIL import Image import io img = Image.open(io.BytesIO(row["source_image"])) ``` ### Quality Scores Every sample has been scored through ScaleEditor's **task-aware quality verification mechanism** across three dimensions, each rated on a 1–3 scale: - **Instruction Following (IF, 1–3):** Does the edited image accurately reflect the intent of the instruction? - **Editing Consistency (EC, 1–3):** Are unedited regions preserved? Is the edit spatially coherent with the source? - **Generation Quality (GQ, 1–3):** Is the output image free of artifacts, distortions, and visual defects? In ScaleEdit, only samples with IF=3, EC≥2, GQ≥2 are retained. ## 🛠️ Highlights ScaleEdit-12M was constructed using the **ScaleEditor** framework, which consists of three stages: 1. **Source Image Expansion** — Curates and expands source images from diverse real and synthetic domains, infusing world knowledge to enable knowledge-grounded editing tasks. 2. **Adaptive Multi-Agent Editing** — An ensemble of specialized agents generates editing instructions and corresponding edited images, adapting strategies per task family. 3. **Task-Aware Quality Verification** — A multi-dimensional scoring system evaluates instruction following, editing consistency, and generation quality, filtering out low-quality samples.  Fine-tuning leading foundation models on ScaleEdit-12M yields consistent improvements: - **Up to +10.4%** on ImgEdit and **+35.1%** on GEdit for general editing benchmarks - **Up to +150.0%** on RISE and **+26.5%** on KRIS-Bench for knowledge-infused editing benchmarks These gains were demonstrated on both UniWorld-V1 and Bagel, showing that open-source agentic pipelines can approach commercial-grade data quality. ## 🌟 Citation ```bibtex @article{chen2026scaleedit, title={ScaleEdit-12M: Scaling Open-Source Image Editing Data Generation via Multi-Agent Framework}, author={Chen, Guanzhou and Cui, Erfei and Tian, Changyao and Yang, Danni and Yang, Ganlin and Qiao, Yu and Li, Hongsheng and Luo, Gen and Zhang, Hongjie}, journal={arXiv preprint arXiv:2603.20644}, year={2026} } @article{tian2026internvl, title={InternVL-U: Democratizing Unified Multimodal Models for Understanding, Reasoning, Generation and Editing}, author={Tian, Changyao and Yang, Danni and Chen, Guanzhou and Cui, Erfei and Wang, Zhaokai and Duan, Yuchen and Yin, Penghao and Chen, Sitao and Yang, Ganlin and Liu, Mingxin and others}, journal={arXiv preprint arXiv:2603.09877}, year={2026} } ```
license: MIT许可证 task_categories: - 图像到图像(image-to-image) language: - 英语 tags: - 图像编辑(image-editing) - 基于指令的编辑(instruction-based-editing) - 多模态(multimodal) - 计算机视觉(computer-vision) - scaleedit - internvl size_categories: - 1000万 < n < 1亿 --- # ScaleEdit-12M: 基于多智能体框架的开源图像编辑数据规模化生成 <div> [](https://arxiv.org/abs/2603.20644) [](https://github.com/gzchen4ai/ScaleEdit-12M) [](https://huggingface.co/datasets/InternVL-U/ScaleEdit-12M) </div> ## 📌 概述 **截至目前规模最大的开源基于指令的图像编辑数据集。** ScaleEdit-12M包含**1240万**经过严格验证的指令-图像配对样本,覆盖真实与合成视觉领域的**23个任务家族**。该数据集由**ScaleEditor**构建,这是一个完全开源的分层多智能体框架,无需使用成本高昂的专有API。  ## 🔥 最新动态 - **[2026/04/03]** 🚀ScaleEdit-12M已在[[Hugging Face]](https://huggingface.co/datasets/InternVL-U/ScaleEdit-12M)发布。 - **[2026/03/24]** 🔥ScaleEdit-12M的论文已在[[arXiv]](https://arxiv.org/abs/2603.20644)上线。 - **[2026/03/06]** 🔥InternVL-U的**技术报告**已发布。可在[[arXiv]](https://arxiv.org/abs/2603.09877)查阅。 ## ✅ 待办事项 - [x] 已发布ScaleEdit-12M数据集 - [ ] 即将发布ScaleEdit-1M子集 - [ ] 即将发布ScaleEditor框架 ## 📊 数据集结构 ### 仓库布局 该数据集按照**23个特定任务子目录**组织,每个子目录包含多个分片的Parquet文件。目录命名遵循`{category_id}_{task_name}`格式: ScaleEdit-12M/ ├── README.md ├── 1.1_style_transfer/ # 全局编辑任务 │ ├── style_transfer_0000.parquet # 每个分片约31.7 GB │ ├── style_transfer_0001.parquet │ ├── ... │ └── style_transfer_0015.parquet ├── 1.2_tone_adjustment/ │ └── tone_adjustment_XXXX.parquet ├── 1.3_viewpoint_transformation/ ├── 1.4_background_replacement/ ├── 2.1_object_addition/ # 对象编辑任务 ├── 2.2_object_removal/ ├── 2.3_object_replacement/ ├── 2.4_action_editing/ ├── 2.5_part_extraction/ ├── 3.1_color_change/ # 属性编辑任务 ├── 3.2_material_change/ ├── 3.3_visual_beautification/ ├── 3.4_count_change/ ├── 3.5_size_change/ ├── 4.1_movie_poster_text_editing/ # 文本编辑任务 ├── 4.2_gui_interface_text_editing/ ├── 4.3_object_surface_text_editing/ ├── 4.4_building_surface_text_editing/ ├── 5.1_perceptual_reasoning/ # 知识注入任务 ├── 5.2_symbolic_reasoning/ ├── 5.3_social_reasoning/ ├── 5.4_scientific_reasoning/ └── 6.1_compositional_editing/ # 组合编辑任务 每个任务文件夹包含**多个Parquet分片**(通常每个约31-32 GB),命名为`{task_name}_{shard_index:04d}.parquet`。分片数量因任务的数据量而异。 ### Parquet文件结构 每个Parquet文件包含以下列: | 列名 | 数据类型 | 描述 | |---|---|---| | `id` | `int64` | 样本唯一标识符 | | `edit_task` | `string` | 任务类别名称(例如 `"style_transfer"`、`"object_addition"`) | | `edit_instruction` | `string` | 自然语言编辑指令 | | `source_image` | `binary` | 源图像(编辑前)的原始二进制字节 | | `edited_image` | `binary` | 编辑后图像的原始二进制字节 | | `source_image_width` | `int64` | 源图像的像素宽度 | | `source_image_height` | `int64` | 源图像的像素高度 | | `edited_image_width` | `int64` | 编辑后图像的像素宽度 | | `edited_image_height` | `int64` | 编辑后图像的像素高度 | | `instruction_following_score` | `int64` | 质量评分:编辑结果遵循指令的程度(1-3分) | | `editing_consistency_score` | `int64` | 质量评分:源图像与编辑后图像的一致性程度(1-3分) | | `generation_quality_score` | `int64` | 质量评分:编辑后图像的整体视觉质量(1-3分) | ### 示例样本行 json { "id": 0, "edit_task": "object_addition", "edit_instruction": "在图像右侧的平台边缘添加一条红白相间的安全护栏。", "source_image": <二进制字节>, "edited_image": <二进制字节>, "source_image_width": 2000, "source_image_height": 1500, "edited_image_width": 2000, "edited_image_height": 1500, "instruction_following_score": 3, "editing_consistency_score": 3, "generation_quality_score": 3 } `source_image`与`edited_image`列以原始二进制字节格式存储图像,可通过以下代码解码为PIL图像: python from PIL import Image import io img = Image.open(io.BytesIO(row["source_image"])) ### 质量评分 所有样本均通过ScaleEditor的**任务感知质量验证机制**从三个维度进行评分,每个维度采用1-3分制: - **指令遵循度(IF, 1-3)**:编辑后的图像是否准确反映了指令的意图? - **编辑一致性(EC, 1-3)**:未编辑区域是否得以保留?编辑操作与源图像的空间对齐是否合理? - **生成质量(GQ, 1-3)**:输出图像是否无伪影、畸变及视觉缺陷? 在ScaleEdit数据集中,仅保留IF=3、EC≥2且GQ≥2的样本。 ## 🛠️ 核心亮点 ScaleEdit-12M基于**ScaleEditor**框架构建,该框架包含三个阶段: 1. **源图像扩展**:从多样的真实与合成视觉领域中精选并扩展源图像,融入世界知识以支持基于知识的编辑任务。 2. **自适应多智能体编辑**:由多个专用智能体组成的集成系统生成编辑指令与对应的编辑后图像,针对不同任务家族适配编辑策略。 3. **任务感知质量验证**:采用多维度评分体系评估指令遵循度、编辑一致性与生成质量,过滤低质量样本。  在ScaleEdit-12M上微调主流基础模型可获得持续的性能提升: - 在通用编辑基准测试ImgEdit上最高提升**+10.4%**,在GEdit上最高提升**+35.1%** - 在知识注入编辑基准测试RISE上最高提升**+150.0%**,在KRIS-Bench上最高提升**+26.5%** 上述性能提升在UniWorld-V1与Bagel两个模型上均得到验证,表明开源智能体流水线可达到商用级别的数据质量。 ## 🌟 引用 bibtex @article{chen2026scaleedit, title={ScaleEdit-12M: 基于多智能体框架的开源图像编辑数据规模化生成}, author={Chen, Guanzhou and Cui, Erfei and Tian, Changyao and Yang, Danni and Yang, Ganlin and Qiao, Yu and Li, Hongsheng and Luo, Gen and Zhang, Hongjie}, journal={arXiv预印本 arXiv:2603.20644}, year={2026} } @article{tian2026internvl, title={InternVL-U: 面向理解、推理、生成与编辑的统一多模态模型民主化}, author={Tian, Changyao and Yang, Danni and Chen, Guanzhou and Cui, Erfei and Wang, Zhaokai and Duan, Yuchen and Yin, Penghao and Chen, Sitao and Yang, Ganlin and Liu, Mingxin and others}, journal={arXiv预印本 arXiv:2603.09877}, year={2026} }




