orpo-text-pairs-full
收藏资源简介:
ORPO文本偏好对(完整版)数据集包含两个版本的偏好对,用于训练语言模型使用ORPO、DPO或类似的基于偏好的对齐方法。数据集包含8,249条经过精炼/过滤的偏好对(推荐使用)和14,214条过滤前的完整数据集。每条记录以JSONL格式存储,包含以下字段:`prompt`(用户对话轮次)、`chosen`(优选回复)、`rejected`(非优选回复)和`meta`(元数据,包括来源数据集、使用的模型和判断信息)。数据集适用于纯文本偏好学习任务,不包含图像。数据来源于多个开源数据集,包括HelpSteer2、MathInstruct、CodeIO-PyEdu-Reasoning和MathV360K,使用时需遵守各来源数据集的许可协议。
The ORPO Text Preference Pairs (Full Version) dataset contains two versions of preference pairs, intended for training language models using ORPO, DPO or similar preference-based alignment methods. The dataset includes 8,249 refined/filtered preference pairs (recommended for use) and 14,214 full unfiltered preference pairs. Each record is stored in JSONL format, with the following fields: `prompt` (user dialogue turn), `chosen` (preferred response), `rejected` (non-preferred response), and `meta` (metadata including source datasets, employed models and judgment information). This dataset is suitable for pure-text preference learning tasks and does not contain any images. It is sourced from multiple open-source datasets including HelpSteer2, MathInstruct, CodeIO-PyEdu-Reasoning and MathV360K, and users must comply with the license agreements of each respective source dataset.




