遇见数据集

ConsistCompose3M

收藏
魔搭社区2026-07-15 更新2026-07-15 收录
官方服务:

资源简介:

![image](./consistcompose3m_banner.webp) <h3 align="center"> ConsistCompose3M: A 3M-Scale Dataset for Unified Multimodal Layout Control in Image Composition </h3> <p align="center"> <a href="https://github.com/OpenSenseNova/ConsistCompose"><img alt="GitHub Stars" src="https://img.shields.io/github/stars/OpenSenseNova/ConsistCompose?style=social"></a> <a href="https://OpenSenseNova.github.io/ConsistCompose/"><img alt="Project Page" src="https://img.shields.io/badge/Project%20Page-ConsistCompose-blue"></a> <a href="https://arxiv.org/abs/2511.18333"><img alt="arXiv" src="https://img.shields.io/badge/arXiv-2511.18333-b31b1b.svg"></a> <a href="https://huggingface.co/sensenova/ConsistCompose"><img src="https://img.shields.io/static/v1?label=%F0%9F%A4%97%20Hugging%20Face&message=Model&color=green"></a> <a href="https://huggingface.co/datasets/sensenova/ConsistCompose3M"><img src="https://img.shields.io/static/v1?label=%F0%9F%A4%97%20Hugging%20Face&message=Dataset&color=yellow"></a> </p> ## **Overview** ConsistCompose3M is a large-scale dataset (~3M samples) dedicated to layout-controllable multi-instance image composition, with significant improvements in scale, diversity, quality and adaptability. It provides millions of diverse multi-instance scenes, identity-preserving samples filtered by CLIP/DINO similarity, and structured spatial-semantic supervision for unified multimodal training. The dataset contains two complementary splits: a layout-grounded text-to-image split built from LayoutSAM with instance-level layout annotations, and a reference-conditioned split for subject-preserving layout-guided generation using high-quality subjects from Subjects200K and UNO. We additionally enrich subject and appearance diversity by integrating virtual try-on data from VITON-HD and DressCode-MR, which are converted into a unified reference-conditioned format. To support comprehensive multimodal learning, we also supplement and standardize general image understanding, text-to-image generation, and image editing data into a consistent annotation schema. All samples are carefully filtered to ensure strong subject consistency and layout alignment. ConsistCompose3M supports both text-only and reference-guided layout-aware generation, and provides a unified, high-quality testbed for research on controllable image composition in multimodal models. We fully open-source this dataset to benefit the research community. ### Key Features - **Scale & Quality**: 3M high-quality paired images with 512×512 / 1024×1024 resolution and multi-aspect ratio support. - **Instance-Level Layout Annotations**: Detailed instance-level spatial annotations for precise layout control. - **Diverse Composition Patterns**: Rich multi-instance layouts and scene combinations for flexible compositional generation. - **Task-Specific Organization**: Well-structured by task type for easy training and evaluation. This dataset enables research and development of: ## Dataset Structure ### Actual Repository Directory Layout ```bash sensenova/ConsistCompose3M/ ├── assets/ # Visual assets (banner, examples) │ └── consistcompose3m_banner.webp ├── jsonl_extended/ # Extended task-specific JSONL annotations │ ├── image_editing/ # Image editing task annotations │ │ ├── Echo_4o_Image_surrel_fantasy_image.jsonl │ │ ├── hq_edit.jsonl │ │ ├── multiedit.jsonl │ │ ├── Nano-consistent-150k.jsonl │ │ ├── omni_edit.jsonl │ │ ├── ShareGPT_4o_edit.jsonl │ │ └── ultra_edit.jsonl │ ├── layout_subject_driven/ # Layout-aware subject-driven generation │ │ ├── Pipeline1_type1.jsonl │ │ ├── Pipeline1_type2.jsonl │ │ └── Pipeline1_type3.jsonl │ ├── layout_t2i/ # Layout-aware text-to-image generation │ │ ├── object365.jsonl │ │ ├── Pipeline1_text2image.jsonl │ │ ├── Pipeline1_type1_text2image.jsonl │ │ ├── Pipeline1_type2_text2image.jsonl │ │ ├── Pipeline1_type3_text2image.jsonl │ │ ├── Pipeline2_text2image.jsonl │ │ └── Pipeline3_text2image.jsonl │ ├── subject_driven/ # Subject-driven generation core data │ │ ├── DressCode-MR_subject_driven.jsonl │ │ ├── Echo_4o_Image_multi_reference_image.jsonl │ │ ├── Pipeline1_subject_driven.jsonl │ │ ├── Pipeline1_type1_subject_driven.jsonl │ │ ├── Pipeline1_type2_subject_driven.jsonl │ │ ├── Pipeline2_subject_driven.jsonl │ │ ├── Pipeline3_subject_driven.jsonl │ │ └── VITON-HD_subject_driven.jsonl │ ├── t2i/ # Text-to-image generation │ │ ├── Echo_4o_Image_instruction_following_image.jsonl │ │ ├── Echo_4o_Image_surrel_fantasy_image.jsonl │ │ ├── text-to-iamge-2M_1024.jsonl │ │ └── text-to-iamge-2M_512.jsonl │ └── understanding/ # Image/text understanding annotations │ ├── Finevision_image_understanding.jsonl │ ├── Finevision_multi_image_understanding.jsonl │ ├── Finevision_text_understanding.jsonl │ └── mammoth_si10M_text_understanding.jsonl ├── DressCode-MR/ # DressCode-MR dataset raw images ├── Pipeline1/ # Pipeline1 generated raw images ├── Pipeline2/ # Pipeline2 generated raw images ├── Pipeline3/ # Pipeline3 generated raw images ├── VITON-HD/ # VITON-HD dataset raw images ├── DressCode-MR.jsonl # DressCode-MR consolidated annotations ├── LayoutSAM.jsonl # LayoutSAM validation annotations ├── Pipeline1.jsonl # Pipeline1 consolidated annotations ├── Pipeline2.jsonl # Pipeline2 consolidated annotations ├── Pipeline3.jsonl # Pipeline3 consolidated annotations ├── VITON-HD.jsonl # VITON-HD consolidated annotations └── README.md # This file ``` ## Citation ```bibtex @InProceedings{Shi_2026_CVPR, author = {Shi, Xuanke and Li, Boxuan and Han, Xiaoyang and Cai, Zhongang and Yang, Lei and Wang, Quan and Lin, Dahua}, title = {ConsistCompose: Unified Multimodal Layout Control for Image Composition}, booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)}, month = {June}, year = {2026}, pages = {495-505} } ```

提供机构:
maas
创建时间:
2026-02-25
二维码
社区交流群
二维码
科研交流群
商业服务