MetaMath-GSM240K
收藏资源简介:
该数据集是`meta-math/MetaMathQA`数据集的一个子集,包含395,000个样本。这个子集是从GSM8K训练集中增强的240,000个样本。数据集包含四个特征:查询(query)、响应(response)、类型(type)和原始问题(original_question)。数据集分为一个训练集(train),包含240,000个样本。
This dataset is a subset of the `meta-math/MetaMathQA` dataset, containing 395,000 samples in total. This subset consists of 240,000 augmented samples sourced from the GSM8K training set. The dataset includes four features: query, response, type, and original_question. It is split into a training set (train) which contains 240,000 samples.
MetaMath-GSM240K 数据集概述
数据集信息
- 许可证: MIT
- 特征:
query: 字符串类型response: 字符串类型type: 字符串类型original_question: 字符串类型
- 分割:
train: 包含240,000个样本,占用238,099,368字节
- 下载大小: 116,355,472字节
- 数据集大小: 238,099,368字节
配置
- 默认配置:
- 数据文件路径:
data/train-*
- 数据文件路径:
数据集来源
- 该数据集是从
meta-math/MetaMathQA数据集中提取的子集,MetaMathQA数据集包含395,000个样本。 - 该子集仅包含从
GSM8K训练集中增强的240,000个样本。
引用
@article{yu2023metamath, title={MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models}, author={Yu, Longhui and Jiang, Weisen and Shi, Han and Yu, Jincheng and Liu, Zhengying and Zhang, Yu and Kwok, James T and Li, Zhenguo and Weller, Adrian and Liu, Weiyang}, journal={arXiv preprint arXiv:2309.12284}, year={2023} }
@article{meng2024pissa, title={Pissa: Principal singular values and singular vectors adaptation of large language models}, author={Meng, Fanxu and Wang, Zhaohui and Zhang, Muhan}, journal={arXiv preprint arXiv:2404.02948}, year={2024} }




