DisciplineGen-1M
收藏资源简介:
DisciplineGen-1M是由上海交通大学等机构联合构建的百万级多学科视觉生成与编辑数据集,旨在支持知识密集型图像的文本到图像生成和图像编辑任务。该数据集包含120万个样本,覆盖数学、物理、化学、生物、地理、计算机科学、经济学、历史、音乐和体育等十个学科,数据来源融合了矢量图形渲染、OCR编辑、程序化合成和大规模过滤等多种技术,以提供结构化的图像-文本对和编辑指令。通过结合可控的语义差异和结构化标注,该数据集旨在解决生成模型在学术视觉内容中因缺乏精确推理和知识约束而导致的可靠性不足问题,推动图像生成从美学合理性向可验证的知识基础视觉创作迈进。
DisciplineGen-1M is a million-scale multi-disciplinary visual generation and editing dataset jointly constructed by Shanghai Jiao Tong University and other institutions. It is designed to support text-to-image generation and image editing tasks for knowledge-intensive images. The dataset contains 1.2 million samples covering ten disciplines including mathematics, physics, chemistry, biology, geography, computer science, economics, history, music and sports. It integrates multiple technologies such as vector graphics rendering, OCR editing, procedural synthesis and large-scale filtering to provide structured image-text pairs and editing instructions. By combining controllable semantic differences and structured annotations, this dataset addresses the problem of insufficient reliability of generative models in academic visual content caused by the lack of precise reasoning and knowledge constraints, and promotes the advancement of image generation from aesthetic plausibility to verifiable knowledge-based visual creation.
数据集名称:DisciplineGen-1M
数据集规模:1.2M 样本(120万样本)
覆盖学科:数学、物理、化学、生物学、地理学、计算机科学、经济学、历史、音乐、体育,共10个学科及其细分子领域。
任务类型:支持文本到图像生成(Text-to-Image Generation)和图像编辑(Image Editing)。
数据集构建方法:采用四种互补的构建管线:
- 基于SVG/TikZ的结构化渲染(Structured Rendering with SVG/TikZ)
- 基于OCR的编辑(OCR-based Editing)
- 大规模文本到图像过滤(Large-scale T2I Filtering)
- 专门化的程序化合成(Specialized Programmatic Synthesis)
数据内容:包含标题(captions)、编辑指令(editing instructions)、结构化注解(structured annotations)以及具有可控语义差异的配对图像。
基准测试表现:在学科相关基准(GenExam、GRADE)上显著优于开源基线模型;在通用推理基准(WISE、RISE)上也展现出更广泛的迁移能力。
发布时间:相关论文发表于 arXiv,编号 arXiv:2607.02290(2026年)。
发布资源:数据集、模型及数据构建管线的源代码将公开发布(包括论文、代码和数据集链接)。
主要贡献机构:上海交通大学、华南理工大学、厦门大学、中国科学技术大学。

- 1DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing上海交通大学; 华南理工大学; 厦门大学; 中国科学技术大学 · 2026年



