Nemotron-Math-Proofs-v2
收藏资源简介:
Nemotron-Math-Proofs-v2是一个数学证明生成、验证和元验证跟踪数据集,旨在支持训练和评估能够生成严格数学证明、并逐步批判或验证证明正确性的模型。数据集包含82,737个样本,涵盖5,752个独特问题,问题来源于nvidia/Nemotron-Math-Proofs-v1数据集中的AoPS(Art of Problems Solving)社区子集。证明、验证跟踪和元验证跟踪是使用DeepSeek-V4-Pro模型在Max推理模式下生成的,遵循DeepSeekMath-V2论文中描述的证明生成和验证提示风格。数据集采用JSONL格式,每个记录包含自然语言证明问题、生成的证明和验证跟踪,主要字段包括:唯一标识符(uuid)、问题陈述(problem)、用于LLM训练的多轮消息序列(messages)、工具定义(tools)、许可证标签(license)、元数据(metadata)、问题来源标签(source)、数据集标签(dataset)、响应类型子集标签(subset,包括证明、验证、元验证)以及下游使用注释字段(used_in)。数据集分为三个子集:证明(24,696个样本)、验证(28,865个样本)和元验证(29,176个样本)。总磁盘大小为15.95 GiB,令牌计数为5,000,839,123。该数据集是先前发布的Nemotron-Math-Proofs-v1和Nemotron-Cascade-2-SFT-Data数据集的扩展,专注于自然语言数学证明数据,但不包含证明细化部分。数据集采用CC BY 4.0许可证,适用于商业和非商业用途,预期用途包括训练LLM进行结构化数学推理和证明生成、生成证明验证跟踪、构建长上下文或多轨迹推理系统以及研究证明有效性、验证器准确性和错误模式。
Nemotron-Math-Proofs-v2 is a mathematical proof generation, verification and meta-verification tracking dataset designed to support the training and evaluation of models capable of generating rigorous mathematical proofs and critically verifying or validating the correctness of proofs step-by-step. It contains 82,737 samples covering 5,752 unique problems, sourced from the AoPS (Art of Problem Solving) community subset of the nvidia/Nemotron-Math-Proofs-v1 dataset. Proofs, verification tracks and meta-verification tracks were generated using the DeepSeek-V4-Pro model in Max inference mode, following the proof generation and verification prompt styles described in the DeepSeekMath-V2 paper. The dataset is stored in JSONL format, with each record containing natural language proof problems, generated proofs and verification tracks. Its main fields include: unique identifier (uuid), problem statement (problem), multi-turn message sequence for LLM training (messages), tool definitions (tools), license tag (license), metadata (metadata), problem source tag (source), dataset tag (dataset), response type subset tag (subset, including proof, verification, meta-verification) and downstream usage annotation field (used_in). The dataset is divided into three subsets: Proof (24,696 samples), Verification (28,865 samples) and Meta-Verification (29,176 samples). It has a total disk size of 15.95 GiB and a token count of 5,000,839,123. This dataset is an extension of the previously released Nemotron-Math-Proofs-v1 and Nemotron-Cascade-2-SFT-Data datasets, focusing on natural language mathematical proof data but excluding the proof refinement section. It is licensed under CC BY 4.0, permitting both commercial and non-commercial use. Its intended applications include training LLMs for structured mathematical reasoning and proof generation, generating proof verification tracks, constructing long-context or multi-trace reasoning systems, and conducting research on proof validity, verifier accuracy and error patterns.
数据集概述
- 名称: Nemotron-Math-Proofs-v2
- 所有者: NVIDIA Corporation
- 创建日期: 2026年5月1日
- 许可证: Creative Commons Attribution 4.0 International License (CC BY 4.0)
- 语言: 英语
- 模态: 文本
- 格式: JSONL
数据集描述
Nemotron-Math-Proofs-v2 是一个数学证明生成、验证和元验证轨迹数据集。问题来源于 nvidia/Nemotron-Math-Proofs-v1 中的 AoPS 子集。数据集包含 82,737 个样本,涵盖 5,752 个独特问题。
数据生成
- 问题来源: 从 nvidia/Nemotron-Math-Proofs-v1 数据集中选取的 AoPS 社区证明类问题。
- 证明与验证轨迹生成: 使用 DeepSeek-V4-Pro 的 Max 推理模式生成证明和验证轨迹,遵循 DeepSeekMath-V2 论文 中描述的提示风格。
- 收集方法: 混合(人工、合成、自动化)
- 标注方法: 混合(人工、合成、自动化)
数据集字段
uuid: 样本的唯一标识符。problem: 问题陈述,源自 nvidia/Nemotron-Math-Proofs-v1 中的 AoPS 子集。messages: 用于 LLM 训练的标准多轮消息序列。tools: 工具定义列表(如有)。license: 每个样本的许可证标签,统一为cc-by-4.0。metadata: 额外的元数据字段。source: 种子问题的来源标签,统一为AoPS。dataset: 数据集/发布标签。subset: 响应类型或子集标签,包括proof、verification或meta-verification。used_in: 保留用于下游使用标注的列表字段。
数据量化
| 分割 | 子集 | 样本数 |
|---|---|---|
| train | proof | 24,696 |
| train | verification | 28,865 |
| train | meta-verification | 29,176 |
| train | 总计 | 82,737 |
- 独特问题数: 5,752
- 总磁盘大小: 15.95 GiB
- Token 数量: 5,000,839,123
预期用途
- 训练 LLM 进行结构化数学推理和证明生成。
- 训练 LLM 生成证明验证轨迹并识别数学论证中的漏洞。
- 构建用于定理证明的长上下文或多轨迹推理系统。
- 研究证明有效性、验证器准确性、错误模式以及自我验证的数学推理。




