FineLeanCorpus
收藏资源简介:
FineLeanCorpus 是一个包含超过 28.5 万个数学问题的数据集,涵盖了广泛的数学领域和难度等级。该数据集经过严格的人类评估,保证了其高正确性。FineLeanCorpus 的创建基于 CriticLean 框架,旨在评估模型区分语义正确和错误形式化的能力。数据集的多样性、难度分布和严格的语义验证,使其成为一个结构平衡的训练环境,有助于解决数学自动形式化中的挑战。
FineLeanCorpus is a dataset containing over 285,000 mathematical problems, covering a wide range of mathematical fields and difficulty levels. This dataset has undergone rigorous human evaluation to ensure its high correctness. FineLeanCorpus is developed based on the CriticLean framework, aiming to evaluate a model's ability to distinguish between semantically correct and incorrect formalizations. The diversity, difficulty distribution, and strict semantic verification of the dataset make it a structurally balanced training environment that helps address the challenges in automated mathematical formalization.
CriticLean 数据集概述
1. 数据集简介
- 名称: CriticLean
- 目的: 通过批评者引导的强化学习框架解决数学陈述形式化问题
- 核心组件:
- CriticLeanGPT: 基于强化学习的模型
- CriticLeanBench: 评估基准
- FineLeanCorpus: 已验证的Lean 4语句数据集
2. 主要数据集
2.1 CriticLeanInstruct
- CriticLean_12K: 全部批评者数据
- CriticLean_4K: 种子数据,用于强化学习
- CriticLean_Mix_48K: 混合数据集,包含12K批评者数据+18K数学数据+18K代码数据
- CriticLean_Mix_16K: 混合数据集,包含4K批评者数据+6K数学数据+6K代码数据
2.2 CriticLeanBench
- 规模: 500对自然语言和Lean 4语句(250正确,250错误)
- 特点:
- 双焦点: 批评能力和Lean 4形式验证
- 多样化覆盖: 不同数学领域和难度级别
- 结构化错误: 常见错误模式设计
- 人工验证: 所有样本经过严格验证
- 评估指标: ACC/TPR/FPR/TNR/FNR
2.3 FineLeanCorpus
- 规模: 285,957条已验证Lean 4语句
- 特点:
- 规模大: 显著超过现有Lean 4问题集
- 高质量: 84%准确率
- 多样化: 覆盖多个数学领域和难度级别
- 严格验证: 通过Lean编译器和CriticLean模型验证
3. 模型变体
| 基础模型 | SFT应用 | SFT数据 | RL应用 | RL数据 | 模型名称 |
|---|---|---|---|---|---|
| Qwen2.5-Instruct | 是 | CriticLean_4K | 否 | * | Qwen2.5-Instruct-SFT(Critic Only) |
| Qwen2.5-Instruct | 是 | CriticLean_Mix_16K | 否 | * | Qwen2.5-Instruct-SFT(16K) |
| Qwen2.5-Instruct | 是 | CriticLean_Mix_48K | 否 | * | Qwen2.5-Instruct-SFT |
| Qwen2.5-Instruct | 是 | CriticLean_Mix_48K | 是 | CriticLean_4K | Qwen2.5-Instruct-SFT-RL |
| Qwen2.5-Instruct | 否 | * | 是 | CriticLean_4K | Qwen2.5-Instruct-RL |
| Qwen3 | 否 | * | 是 | CriticLean_4K | Qwen3-RL |
4. 数据来源
- AOPs
- DeepMath-103k
- NuminaMath-TIR
- DeepTheorem
- DeepScaleR
- DAPO-Math-17k
- Omni-Math
- InwqMath
- BlueMO
- TAL-SCQ5K
- OnlineMathContest
- Multi-Source Math Competition
5. 许可信息
- 许可证: Apache-2.0 license
6. 引用信息
bibtex @misc{peng2025criticleancriticguidedreinforcementlearning, title={CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization}, author={Zhongyuan Peng and Yifan Yao and Kaijing Ma and Shuyue Guo and Yizhe Li and Yichi Zhang and Chenchen Zhang and Yifan Zhang and Zhouliang Yu and Luming Li and Minghao Liu and Yihang Xia and Jiawei Shen and Yuchen Wu and Yixin Cao and Zhaoxiang Zhang and Wenhao Huang and Jiaheng Liu and Ge Zhang}, year={2025}, eprint={2507.06181}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2507.06181}, }

- 1CriticLean: Critic-Guided Reinforcement Learning for Mathematical Formalization字节跳动种子南京大学 · 2025年



