davincimath/davinci-math
收藏资源简介:
daVinci-Math是一个统一的、多阶段的数学推理数据集。与为中期训练、监督微调和强化学习单独策划资源不同,它构建了一个单一的、阶段感知的管道,从公共数学解题资源开始,应用统一的清洗和去重,并将每个问题路由到最有效的训练阶段。该数据集旨在支持三个训练阶段:中期训练(提供广泛的数学覆盖和多样化的推理模式)、监督微调(提供高质量的后训练问题,带有验证过的推理轨迹)和强化学习(提供一个较小的、具有挑战性的、规则可验证的子集,用于基于奖励的优化)。
daVinci-Math is a unified, multi-stage mathematical reasoning dataset. Unlike separately curated resources for pre-training, supervised fine-tuning, and reinforcement learning, it constructs a single, stage-aware pipeline that starts from public mathematical problem-solving resources, applies unified cleaning and deduplication, and routes each problem to the most effective training stage. This dataset aims to support three training stages: pre-training, which provides extensive mathematical coverage and diverse reasoning patterns; supervised fine-tuning, which provides high-quality post-training problems with verified reasoning trajectories; and reinforcement learning, which provides a smaller, challenging, rule-verifiable subset for reward-based optimization.



