orca dpo pairs
收藏资源简介:
orca dpo pairs数据集由韩国首尔延世大学等机构创建,包含约13,000条提示、选择响应和拒绝响应的配对,主要用于指令翻译、写作、常识和数学推理等任务。数据集通过GPT-3.5-turbo模型进行增强,生成了37,000条新的提示,并使用预训练的RM-Gemma-7B模型对响应进行评分和排序。该数据集主要用于优化大型语言模型的偏好学习,旨在提高模型在执行人类指令时的准确性和安全性。
The Orca DPO Pairs dataset was developed by institutions including Yonsei University in Seoul, Republic of Korea, and contains approximately 13,000 prompt, chosen response, and rejected response pairs, primarily used for tasks such as instruction translation, writing, common sense reasoning and mathematical reasoning. The dataset was augmented using the GPT-3.5-turbo model to generate 37,000 new prompts, and responses were scored and ranked with the pre-trained RM-Gemma-7B model. This dataset is mainly employed to optimize preference learning for large language models, with the goal of improving the accuracy and safety of models when executing human instructions.




