遇见数据集

AI-MO/B2-UniMath

收藏
Hugging Face2026-05-13 更新2026-06-14 收录
官方服务:

资源简介:

B2-UniMath是一个大型多语言数学数据集,包含从数学PDF中提取的问题-解决方案-答案对。该数据集专为训练和研究自然语言数学推理而设计,特别侧重于大学水平的内容。数据来源于CommonCrawl档案中的PDF,经过重新获取、OCR识别,并通过多个过滤和处理阶段生成。数据集包含约400万个问答对,格式包括问题陈述、可选选项、解决方案、最终答案以及辅助注释。每个样本都提供了丰富的元数据,如唯一标识符、来源文档信息、难度评分(涵盖小学到研究水平)、语言检测、问题类型和有效性标签。此外,每个样本都使用Qwen3-32B模型将原始语言翻译成英文。数据集主要用于数学文本的持续预训练、数学问答数据的监督微调、数学语料库的过滤和消融分析,以及数学文本的翻译任务。数据集中大学水平内容占主导(约80.2%为大学水平,19%为大学竞赛水平),语言分布广泛,其中英语占比最高(约60.5%)。

B2-UniMath is a large multilingual mathematics dataset of question-solution-answer pairs extracted from mathematical PDFs. The dataset is designed for training and studying natural-language mathematical reasoning, with a strong emphasis on university-level content. The data is collected from PDF data contained in the CommonCrawl archives, which are refetched, OCRed, and passed through several filtering and processing stages. It contains approximately 4 million Q&A pairs, with each sample including problem statement, optional choices, solution, final answer, and auxiliary annotations. Rich metadata is provided for each sample, such as unique identifier, source document information, difficulty scores (covering elementary to research levels), language detection, question type, and validity tags. Additionally, each sample contains a translation from the base language into English using Qwen3-32B as the translator. The dataset is intended for continued pretraining on mathematical text, supervised fine-tuning on math Q&A data, filtering, ablation, and multilingual analysis of math corpora, and language-specific training tasks like translation of mathematical texts. It is predominantly university-level (about 80.2% university, 19% university competition), with a wide language distribution where English accounts for approximately 60.5%.

提供机构:
AI-MO
二维码
社区交流群
二维码
科研交流群
商业服务