ETCHR-GRPO-10K
收藏资源简介:
# ETCHR GRPO-10K 📖<a href="https://arxiv.org/abs/2605.23897">Paper</a> | 🏠<a href="https://github.com/InternLM/ETCHR">Homepage</a > | 🤗<a href="https://huggingface.co/internlm/ETCHR-FLUX.2-klein-9B">ETCHR-FLUX.2-klein-9B Model</a > | 🤗<a href="https://huggingface.co/datasets/BeichenZhang/ETCHR-SFT-400K">ETCHR SFT-400K Dataset</a > | 🤗<a href="https://huggingface.co/datasets/internlm/ETCHR-GRPO-10K">ETCHR GRPO-10K Dataset</a > | 🤗<a href="https://huggingface.co/datasets/internlm/DL3DV-2k">DL3DV-2K Benchmark</a > ETCHR GRPO-10K is the GRPO training data for further enhance ETCHR's editing capaibility in assisting understanding models. It contains 10000 samples of five tasks (Fine-grained Perception, Chart Understanding, Maze Solving, Jigsaw Puzzle and Spatial Understanding). Each sample contains the image to be edited, an editing instruction, and a corresponding understanding task associated with this image for measuring the editing quality via guidance reward. ## 📢 News - 🚀 [2026/05/22] We have released the training and evaluation code of ETCHR. - 🚀 [2026/05/21] We have released the [ETCHR-FLUX.2-klein-9B Model](https://huggingface.co/internlm/ETCHR-FLUX.2-klein-9B), [ETCHR-SFT-400K Dataset](https://huggingface.co/datasets/BeichenZhang/ETCHR-SFT-400K) and [ETCHR GRPO-10K Dataset](https://huggingface.co/datasets/internlm/ETCHR-GRPO-10K). ## 🌈 Overview We are thrilled to introduce ETCHR (Editing To Clarify and Harness Reasoning), a novel question-conditioned, reasoning-aware image editor designed to serve as a decoupled visual reasoning assistant for Multimodal Large Language Models (MLLMs). By decoupling the specialized image editor from the downstream understanding model, ETCHR bridges the critical bottleneck where a purely textual chain of thought fails in fine-grained focus or complex spatial transformations. </p> <p style="text-align: center;"> <img src="assets/overview.png" alt="Teaser" width="100%"> </p> ## 💡 Highlights - 🔥 **Decoupled & Plug-and-Play:** ETCHR functions as a separate module, allowing it to assist diverse downstream MLLMs (such as Qwen3-VL-8B, Gemini-3.1-Flash-Lite, or Kimi K2.5) without requiring any task-specific fine-tuning on the understanding models themselves. - 🔥 **Naturally Reflective Pipeline:** Introduces an Edit-Verify-Reason inference mechanism where the understanding model filters out noisy or flawed edits, reverting safely to the original image when verification fails. ## 🛠️ Usage You can find all source images, edit instructions and qa list for Guidance Reward in `GRPO-10K.parquet`. See [https://github.com/InternLM/ETCHR/blob/master/RL/RL.md](https://github.com/InternLM/ETCHR/blob/master/RL/RL.md) for further details. ## ✒️Citation If you find this project useful, please kindly cite: ``` @article{zhang2026etchr, title={ETCHR: Editing To Clarify and Harness Reasoning}, author={Beichen Zhang, Yuhong Liu, Jinsong Li, Yuhang Zang, Jiaqi Wang, Dahua Lin}, journal={arXiv preprint arXiv:2605.23897}, year={2026} } ```



