OpenMath-200k
收藏资源简介:
OpenMath-200k是一个高质量、大规模的数学推理数据集,包含约20万个数学问题及其逐步解答。该数据集旨在支持数学推理模型的训练与评估,特别是针对思维链推理能力的提升。核心由两个子集组成:推理子集(约10.2万个样本,解答带有显式的思维过程标签,引导结构化逻辑推理)和标准子集(约9.8万个样本,提供普通的思维链解答作为基准训练数据)。每个样本包含唯一标识符、问题陈述、详细逐步解答、最终答案、所属数学主题(如代数、几何、微积分等)、难度等级(易、中、难)、解答是否经过格式验证的标志以及推理格式类型。数据覆盖广泛的数学领域,包括代数、几何、三角学、微积分、数论、概率、统计和组合数学等,难度分布均衡(难题50%、中等难度35%、简单题15%)。数据已按90%训练、5%验证、5%测试的比例划分,所有解答经过质量验证,格式纯净,专注于问题与解答本身。适用于训练和微调大型语言模型进行数学问题求解、提升思维链推理能力以及作为评估模型数学推理性能的基准。
OpenMath-200k is a high-quality, large-scale mathematical reasoning dataset containing approximately 200,000 math problems and their step-by-step solutions. This dataset is designed to support the training and evaluation of mathematical reasoning models, particularly for enhancing chain-of-thought reasoning abilities. It consists of two core subsets: the reasoning subset (about 102,000 samples, whose solutions are tagged with explicit thinking process labels to guide structured logical reasoning) and the standard subset (about 98,000 samples, which provide conventional chain-of-thought solutions as benchmark training data). Each sample includes a unique identifier, problem statement, detailed step-by-step solution, final answer, affiliated mathematical topic (e.g., algebra, geometry, calculus, etc.), difficulty level (easy, medium, hard), a flag indicating whether the solution has passed format validation, and the reasoning format type. The dataset covers a wide range of mathematical domains including algebra, geometry, trigonometry, calculus, number theory, probability, statistics, combinatorics and more, with a balanced difficulty distribution: 50% hard problems, 35% medium problems, and 15% easy problems. The data is split at a ratio of 90% training, 5% validation, and 5% testing. All solutions have undergone quality verification and feature pure formats, focusing solely on the problems and their corresponding solutions. It is applicable for training and fine-tuning large language models for mathematical problem-solving, improving chain-of-thought reasoning capabilities, and serving as a benchmark for evaluating the mathematical reasoning performance of models.




