遇见数据集

MARIO-Math-Reasoning/AlphaMath-Trainset

收藏
Hugging Face2024-06-20 更新2024-06-25 收录
官方服务:

资源简介:

--- # For reference on dataset card metadata, see the spec: https://github.com/huggingface/hub-docs/blob/main/datasetcard.md?plain=1 # Doc / guide: https://huggingface.co/docs/hub/datasets-cards {} --- # Dataset Card for AlphaMath Almost Zero <!-- Provide a quick summary of the dataset. --> This is the round 3 training data for [AlphaMath Almost Zero: Process Supervision Without Process](https://arxiv.org/abs/2405.03553). The solution process was automatically generated by the model in round 2, without GPT or Human annotations. ## Dataset Details 1. The question-answer pairs are extracted from the train split of [GSM8k](https://huggingface.co/datasets/openai/gsm8k) and [MATH](https://github.com/hendrycks/math). 2. Both positive and negative examples are included, for training both policy and value models.

这是AlphaMath Almost Zero项目的第三轮训练数据,其解题过程是由模型在第二轮自动生成的,没有使用GPT或人工标注。数据集中的问答对是从GSM8k和MATH数据集的训练分割中提取的,并且包含了正例和负例,用于训练策略模型和价值模型。

提供机构:
MARIO-Math-Reasoning
原始信息汇总

数据集卡片 AlphaMath Almost Zero

概述

这是AlphaMath Almost Zero: Process Supervision Without Process的第三轮训练数据。解决方案过程由第二轮模型自动生成,没有使用GPT或人工注释。

数据集详情

  1. 问题-答案对从GSM8kMATH的训练集中提取。
  2. 包含正例和负例,用于训练策略和价值模型。
搜集汇总
背景与挑战
背景概述
该数据集是AlphaMath Almost Zero项目的第三轮训练数据,其解题过程由模型自动生成,无需人工或GPT标注。它从GSM8k和MATH数据集的训练分割中提取问答对,包含正例和负例,用于训练策略模型和价值模型。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务