selfcorrexp2/llama3_sft_less_corr_train_on_corr_dpo_gen1_math
收藏数据链接:
官方服务:
资源简介:
该数据集是一个包含索引、提示、答案序列、正确答案和首次奖励布尔值等字段的数据集,用于训练模型。数据集分为训练集,共有7496个样本,数据集大小为146184875字节。
This dataset includes fields such as index, prompt, answer sequence, correct answer, and first reward boolean, which is used for training models. The dataset is divided into a training set with a total of 7496 samples and a dataset size of 146184875 bytes.
提供机构:
selfcorrexp2


