MINT-CoT
收藏资源简介:
MINT-CoT数据集由香港中文大学多媒体实验室构建,包含54,000个数学问题,每个问题都与其推理步骤在token级别上与视觉区域对齐,并伴随一个严格的数据生成流程。该数据集旨在解决现有方法在解决数学问题时所面临的三个主要限制:依赖于粗粒度的框状图像区域、视觉编码器对数学内容的感知有限以及依赖于外部能力进行视觉修改。MINT-CoT数据集通过在推理步骤中自适应地交织相关视觉token,为训练多模态数学推理模型提供了基础。
The MINT-CoT dataset was constructed by the Multimedia Laboratory of The Chinese University of Hong Kong. It contains 54,000 mathematical problems, each of which has its reasoning steps aligned with corresponding visual regions at the token level and is accompanied by a rigorous data generation pipeline. This dataset aims to resolve three major limitations faced by existing approaches to mathematical problem solving: reliance on coarse-grained bounding-box image regions, limited perception of mathematical content by visual encoders, and dependence on external capabilities for visual modification. The MINT-CoT dataset provides a solid foundation for training multimodal mathematical reasoning models by adaptively interleaving relevant visual tokens within reasoning steps.
MINT-CoT数据集概述
基本信息
- 数据集名称: MINT-CoT
- 官方论文: MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning
- 数据集地址: HuggingFace数据集
- 模型地址: HuggingFace模型
- 发布时间: 2025年6月6日
数据集简介
- 核心目标: 解决多模态领域数学推理中Chain-of-Thought (CoT) 的扩展挑战
- 创新点:
- 提出MINT-CoT方法,通过Interleave Token动态选择数学图形中的任意形状视觉区域
- 将视觉标记自适应地交织到文本推理步骤中
数据集内容
- 数据规模: 包含54K数学问题
- 数据特点:
- 每个推理步骤与标记级别的视觉区域对齐
- 包含严格的数据生成流程
训练策略
- 第一阶段: 纯文本CoT监督微调 (SFT)
- 第二阶段: 交织CoT监督微调 (SFT)
- 第三阶段: 交织CoT强化学习 (RL)
评估方法
- 评估工具: VLMEvalKit
- 评估基准: MathVista_MINI
- 评估模型: Qwen2-VL-7B-Instruct
依赖工具
- 基础框架: R1-V
- 训练框架: LLaMA-Factory
- 辅助工具: Mulberry

- 1MINT-CoT: Enabling Interleaved Visual Tokens in Mathematical Chain-of-Thought Reasoning香港中文大学多媒体实验室(CUHK MMLab) · 2025年



