selfcorrexp2/llama3_sft_ift_morecorr_more_datatmp10_vllmexp
收藏数据链接:
官方服务:
资源简介:
该数据集包含多个字段,用于存储索引、提示文本、奖励标识、答案序列、真实标签、代理标签和二次奖励标识。数据集被划分为训练集,共有30000个示例,文件大小为85574684字节。不过,数据集的具体内容和用途在README文件中并未描述。
The dataset includes multiple fields for storing index, prompt text, reward indicators, answer sequences, ground truth labels, proxy labels, and second reward indicators. The dataset is split into a training set with a total of 30,000 examples, and the file size is 85574684 bytes. However, the specific content and purpose of the dataset are not described in the README file.
提供机构:
selfcorrexp2


