DisciplineGen-1M
收藏资源简介:
DisciplineGen-1M是一个百万规模的多学科数据集,旨在支持文本到图像(T2I)生成和图像编辑任务。该数据集解决了现有图像生成模型的关键空白:虽然这些模型可以生成视觉上吸引人的自然图像,但在生成依赖于学科概念、符号结构和精确空间关系的知识密集型图表时仍不可靠。数据集包含120万个样本,覆盖数学、物理、化学、生物、地理、计算机科学、经济学、历史、音乐、体育等10多个学科,提供标题、编辑指令、结构化注释以及具有可控语义差异的配对图像。
DisciplineGen-1M is a million-scale multidisciplinary dataset designed to support text-to-image (T2I) generation and image editing tasks. This dataset addresses a critical gap in existing image generation models: while these models can produce visually appealing natural images, they remain unreliable when generating knowledge-intensive diagrams that rely on disciplinary concepts, symbolic structures and precise spatial relationships. The dataset contains 1.2 million samples, covering over 10 disciplines including mathematics, physics, chemistry, biology, geography, computer science, economics, history, music and sports, and provides captions, editing instructions, structured annotations as well as paired images with controllable semantic differences.
数据集概述:DisciplineGen-1M
DisciplineGen-1M 是一个百万级的多学科数据集,旨在支持文本到图像(T2I)生成和图像编辑任务。它解决了现有图像生成模型在知识密集型图表生成上的不足,这类图表依赖于学科概念、符号结构和精确空间关系。
主要特点
- 双任务支持:同时支持文本生成图像和图像编辑。
- 学科知识推理:引入了基于学科知识的推理-生成模型。
- 大规模样本:包含约 120 万样本。
- 多学科覆盖:涵盖数学、物理、化学、生物学、地理学、计算机科学、经济学、历史、音乐、体育等 10 余个学科。
- 结构化标注:包含标题、编辑指令、结构化标注以及具有可控语义差异的配对图像。
数据集构建
采用可扩展框架,结合了四种互补方法:
- 矢量图形渲染(SVG/TikZ):从矢量图形格式进行结构化渲染。
- 基于OCR的编辑:利用光学字符识别创建编辑配对。
- 大规模T2I过滤:大规模过滤文本到图像数据。
- 专用程序化合成:精选的程序化生成学科内容。
这些流程生成了标题、编辑指令、结构化标注以及具有可控语义差异的配对图像。
数据集示例特点
- 长且信息密集的提示词。
- 跨细粒度子领域的多样化主题覆盖。
- 多种图像类别。
- 多样的分辨率和宽高比。
研究成果
该方法在学科相关基准(GenExam、GRADE)上相比开源基线取得了实质性改进,并在通用推理知情基准(WISE、RISE)上展现出更广泛的迁移能力。这证明了大规模结构化学术视觉数据对于将图像生成从美学合理性转向可验证的知识驱动视觉创造至关重要。
引用
如您的研究使用该工作,请引用:
bibtex @misc{wang2026disciplinegen1mlargescaledatasetmultidisciplinary, title={DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing}, author={Zhaokai Wang and Mingxin Liu and Zirun Zhu and Ziqian Fan and Yiguo He and Mohan Zhang and Leyao Gu and Xiangyu Zhao and Ning Liao and Shaofeng Zhang and Xuanhe Zhou and Zhihang Zhong and Junchi Yan and Xue Yang}, year={2026}, eprint={2607.02290}, archivePrefix={arXiv}, primaryClass={cs.CV}, url={https://arxiv.org/abs/2607.02290}, }





