jescy525/archon-sft-v1-dpo
收藏资源简介:
这是一个用于ARCHON对齐训练的DPO(直接偏好优化)偏好对数据集,包含131098个数据对。数据格式与trl.DPOTrainer兼容,每个样本包括提示(prompt)、优选回复(chosen)和拒绝回复(rejected),其中chosen和rejected以ChatML格式组织,包含系统、用户和助手角色内容。数据集来源于多个公开数据集,包括argilla/distilabel-intel-orca-dpo-pairs、Intel/orca_dpo_pairs、HuggingFaceH4/ultrafeedback_binarized和mlabonne/orpo-dpo-mix-40k,并经过采样确保100%为英语内容,任务类型均为dpo_pair。该数据集旨在支持模型通过直接偏好优化进行对齐训练,是ARCHON SFT v1的实际数据,而非规划文档。
This is a DPO (Direct Preference Optimization) preference pairs dataset for ARCHON alignment training, containing 131,098 pairs. The data format is compatible with trl.DPOTrainer, with each sample including a prompt, chosen response, and rejected response, where chosen and rejected are organized in ChatML format with system, user, and assistant roles. The dataset is sourced from multiple public datasets including argilla/distilabel-intel-orca-dpo-pairs, Intel/orca_dpo_pairs, HuggingFaceH4/ultrafeedback_binarized, and mlabonne/orpo-dpo-mix-40k, and is sampled to ensure 100% English content with all task types as dpo_pair. It is designed to support model alignment training via direct preference optimization and represents the actual data for ARCHON SFT v1, not planning documents.




