MARIO-Math-Reasoning/AlphaMath-Trainset
收藏资源简介:
--- # For reference on dataset card metadata, see the spec: https://github.com/huggingface/hub-docs/blob/main/datasetcard.md?plain=1 # Doc / guide: https://huggingface.co/docs/hub/datasets-cards {} --- # Dataset Card for AlphaMath Almost Zero <!-- Provide a quick summary of the dataset. --> This is the round 3 training data for [AlphaMath Almost Zero: Process Supervision Without Process](https://arxiv.org/abs/2405.03553). The solution process was automatically generated by the model in round 2, without GPT or Human annotations. ## Dataset Details 1. The question-answer pairs are extracted from the train split of [GSM8k](https://huggingface.co/datasets/openai/gsm8k) and [MATH](https://github.com/hendrycks/math). 2. Both positive and negative examples are included, for training both policy and value models.
这是AlphaMath Almost Zero项目的第三轮训练数据,其解题过程是由模型在第二轮自动生成的,没有使用GPT或人工标注。数据集中的问答对是从GSM8k和MATH数据集的训练分割中提取的,并且包含了正例和负例,用于训练策略模型和价值模型。
数据集卡片 AlphaMath Almost Zero
概述
这是AlphaMath Almost Zero: Process Supervision Without Process的第三轮训练数据。解决方案过程由第二轮模型自动生成,没有使用GPT或人工注释。



