SciIR-82k
收藏资源简介:
SciIR-82k是由华中科技大学等机构构建的大规模科学图像推理数据集,旨在解决文本到图像生成模型在科学图像领域语义对齐与逻辑推理的不足。该数据集包含超过82,000个高质量的科学图文对,数据源自《自然》及《自然通讯》期刊的CC BY 4.0许可文章,经过自动化布局分析和两阶段筛选流程提取而成。数据集构建过程基于皮尔斯符号学三元论,将科学推理形式化为实体结构、科学过程与科学定律三个核心维度,并辅以科学推理思维链(Sci-RCoT)注释来显式建模底层视觉逻辑。其核心应用在于为科学图像生成提供过程导向的训练与评估资源,以提升模型在编码物理定律、准确拓扑及因果逻辑等方面的严谨性。
SciIR-82k is a large-scale scientific image reasoning dataset developed by Huazhong University of Science and Technology and other institutions, aiming to address the shortcomings of text-to-image generation models in terms of semantic alignment and logical reasoning in the scientific image domain. This dataset contains over 82,000 high-quality scientific image-text pairs, extracted from CC BY 4.0 licensed articles published in *Nature* and *Nature Communications* via automated layout analysis and a two-stage screening pipeline. The dataset construction process is based on Peirce's triadic theory of semiotics, which formalizes scientific reasoning into three core dimensions: entity structure, scientific processes, and scientific laws, and is supplemented with Scientific Reasoning Chain of Thought (Sci-RCoT) annotations to explicitly model the underlying visual logic. Its core application lies in providing process-oriented training and evaluation resources for scientific image generation, to enhance the rigor of models in encoding physical laws, accurate topology, causal logic and other relevant aspects.

- 1SciIR: A Large-scale Training Dataset and Benchmark for Scientific Image Reasoning Generation华中科技大学·计算机科学与技术学院; 山东大学·空天科学与工程学院; 清华大学·电子工程系 · 2026年



