遇见数据集

upb-nlp/Dolci-Instruct-DPO-en

收藏
Hugging Face2026-05-12 更新2026-05-31 收录
官方服务:

资源简介:

该数据集是一个偏好对数据集,用于训练和评估语言模型的偏好对齐。它包含chosen(优选响应)和rejected(被拒绝响应)两个主要特征,每个响应包括内容、国家、语言、用户代理、毒性标记等字段。数据集还记录了生成响应的模型名称、提示ID和偏好类型。数据集规模为208,869个训练示例,总大小约1.04 GB,适用于强化学习从人类反馈(RLHF)或直接偏好优化(DPO)等任务。

This dataset is a preference pair dataset designed for training and evaluating language model preference alignment. It includes two main features: chosen (preferred responses) and rejected (rejected responses), each with fields such as content, country, language, user agent, toxicity flags, etc. The dataset also records the model names that generated the responses, prompt IDs, and preference types. It consists of 208,869 training examples with a total size of approximately 1.04 GB, suitable for tasks like reinforcement learning from human feedback (RLHF) or direct preference optimization (DPO).

提供机构:
upb-nlp
二维码
社区交流群
二维码
科研交流群
商业服务