selfcorrexp2/llama3_sft_ift_morecorr_more_datatmp07_vllmexp
收藏数据链接:
官方服务:
资源简介:
该数据集包含了七个字段:索引(idx),提示文本(prompt),奖励布尔序列(rewards),答案字符串序列(answers),真实标签(gt),代理标签(proxy_label),第二个奖励布尔序列(second_rewards)。数据集分为训练集(train),共有30000个样本。数据集的下载大小为27917620字节,总大小为78541002字节。
The dataset includes seven fields: index (idx), prompt text (prompt), reward boolean sequence (rewards), answer string sequence (answers), ground truth label (gt), proxy label (proxy_label), and second reward boolean sequence (second_rewards). The dataset is split into a training set (train) with a total of 30,000 samples. The download size of the dataset is 27,917,620 bytes, and the total size is 78,541,002 bytes.
提供机构:
selfcorrexp2


