xinlai/Math-Step-DPO-10K
收藏资源简介:
Math-Step-DPO-10K是一个高质量的逐步偏好数据集,用于数学推理。该数据集是Step-DPO项目的一部分,Step-DPO是一种简单、有效且数据高效的方法,用于提升大语言模型的数学推理能力。该数据集包含多个字段,如数据集名称、提示、初始推理步骤、选择、拒绝、完整选择、完整拒绝和答案等。数据集的应用效果显著,例如在Qwen2-72B-Instruct模型上,Step-DPO在MATH和GSM8K测试集上分别取得了70.8%和94.0%的分数,超过了包括GPT-4-1106、Claude-3-Opus和Gemini-1.5-Pro在内的一系列闭源模型。
Math-Step-DPO-10K is a high-quality step-by-step preference dataset for mathematical reasoning. It is part of the Step-DPO project, which is a simple, effective and data-efficient approach to enhancing the mathematical reasoning capabilities of large language models (LLMs). The dataset contains multiple fields such as dataset name, prompt, initial reasoning steps, chosen, rejected, full chosen, full rejected and answer. It has achieved notable application effects: for example, on the Qwen2-72B-Instruct model, Step-DPO achieved scores of 70.8% and 94.0% on the MATH and GSM8K test sets respectively, outperforming a series of closed-source models including GPT-4-1106, Claude-3-Opus and Gemini-1.5-Pro.
Math-Step-DPO-10K 数据集概述
数据集信息
- 语言: 英语
- 特征:
dataset: 字符串类型prompt: 字符串类型initial_reason_steps: 字符串类型chosen: 字符串类型rejected: 字符串类型full_chosen: 字符串类型full_rejected: 字符串类型answer: 字符串类型
- 分割:
train: 包含 10795 个样本,占用 26528471 字节
- 下载大小: 11985248 字节
- 数据集大小: 26528471 字节
配置
- 配置名称:
default - 数据文件:
train: 路径为data/train-*
标签
dpo




