遇见数据集

AoPS-Instruct

收藏
arXiv2025-09-30 收录
数据链接:
官方服务:

资源简介:

该数据集包含了从“艺术解决问题”论坛中提取的超过60万对高质量的问题与答案,旨在提升大型语言模型在推理能力上的表现。此外,该数据集中还包含了从常用的数学基准测试中清洗过的数据,以避免与现有数据重叠。规模上,该数据集超过了60万对问题与答案,任务则是针对奥林匹克级别的数学问题,对大型语言模型进行训练和评估。

This dataset contains over 600,000 high-quality question-answer pairs extracted from the Art of Problem Solving (AoPS) forum, with the objective of enhancing the reasoning capabilities of large language models (LLMs). Additionally, it includes curated data cleaned from common mathematical benchmarks to avoid overlap with existing datasets. In terms of scale, the dataset has more than 600,000 question-answer pairs, and the tasks are designed to train and evaluate LLMs on Olympiad-level mathematical problems.

搜集汇总
数据集介绍
AoPS-Instruct 数据集图片
背景与挑战
背景概述
AoPS-Instruct是一个基于在线奥林匹克级别数学问题的数据集,专为语言模型(LLMs)的训练和抗污染评估设计。它包含从AoPS平台爬取的数据,经过解析和重写处理成训练格式,并通过第三方在Hugging Face上提供。数据集还附带了完整的处理代码和评估工具,以支持研究和实验。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务