mlfoundations-dev/multiple_samples_none_seed_code
收藏资源简介:
该数据集包含多个字段,其中有问题(problem)、70B蒸馏模型的第一响应(r1_distill_70b_response)、原始行索引(__original_row_idx)、多数响应(_majority_responses)、经过验证的70B蒸馏模型响应(verified_r1_distill_70b_response)以及对话信息(conversations,包括对话来源和内容)。数据集被划分为训练集(train),包含7500个示例,大小为467122274字节。
The dataset includes multiple fields such as the problem, the first response from the 70B distillation model (r1_distill_70b_response), the original row index (__original_row_idx), the majority response (_majority_responses), the verified response from the 70B distillation model (verified_r1_distill_70b_response), and conversation information (conversations, including the source and content of the conversation). The dataset is split into a training set (train) with 7500 examples, totaling 467122274 bytes in size.



