HUNYUANPROVER数据集
收藏资源简介:
HUNYUANPROVER数据集由腾讯混元团队创建,旨在解决自动定理证明中的数据稀疏问题。该数据集包含3万条合成实例,每条实例包括自然语言中的原始问题、通过自动形式化转换的陈述以及由HunyuanProver生成的证明。数据集通过开源数学问题和自然语言数学问题生成,经过多次迭代优化,最终用于训练和改进自动定理证明模型。该数据集的应用领域主要集中在自动定理证明,旨在提升模型在复杂数学问题上的推理和证明能力。
The HUNYUANPROVER dataset was developed by the Tencent Hunyuan Team to tackle the problem of data sparsity in automated theorem proving. It contains 30,000 synthetic instances, each including the original natural language problem, the formal statement converted via automated formalization, and the proof generated by HunyuanProver. The dataset is generated from open-source mathematical problems and natural language mathematical problems, and has undergone multiple iterative optimizations before being ultimately used for training and improving automated theorem proving models. The primary application domain of this dataset is automated theorem proving, aiming to enhance the model's reasoning and proof capabilities for complex mathematical problems.




