遇见数据集

friederrr/PolyUniMath

收藏
Hugging Face2026-05-26 更新2026-05-31 收录
官方服务:

资源简介:

PolyUniMath 是一个从数学PDF中提取的大规模多语言数学问题-解答-答案对数据集。该数据集专为训练和研究自然语言数学推理而设计,特别强调大学水平的内容。数据集包含约400万个问答对,主要格式包括问题陈述、可选选项、解答过程、最终答案以及辅助注释。每个样本都提供了丰富的元数据,这些元数据来自数据生成管道的各个阶段,并且使用Qwen3-32B作为翻译器,将基础语言翻译成英语。数据来源于CommonCrawl档案中的PDF数据,经过重新获取、OCR处理,并通过多个过滤和处理阶段生成完整的管道。

PolyUniMath is a large multilingual mathematics dataset of question-solution-answer pairs extracted from mathematical PDFs. The dataset is designed for training and studying natural-language mathematical reasoning, with a strong emphasis on university-level content. Sample count: approximately 4 million Q&A pairs; Main focus: university-level mathematics; Format: problem, optional choices, solution, final answer, and auxiliary annotations; Metadata: We provide rich metadata for each sample collected from various stages of our data generation pipeline; Translation: Each sample contains a translation from the base language into English using Qwen3-32B as the translator. The data itself is collected from PDF data contained in the CommonCrawl archives, which we refetch, OCR, and then pass through several filtering and processing stages.

提供机构:
friederrr
二维码
社区交流群
二维码
科研交流群
商业服务