遇见数据集

ZQW329/hh-rlhf

收藏
Hugging Face2026-05-17 更新2026-05-31 收录
官方服务:

资源简介:

该数据集包含两种类型的数据:1. 关于帮助性和无害性的人类偏好数据,来自论文《Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback》,用于训练偏好(或奖励)模型以进行后续的RLHF训练,但不适用于监督训练对话代理;2. 人类生成和标注的红队对话数据,来自论文《Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned》,用于理解红队攻击模型的方式和成功/失败类型,不适用于微调或偏好建模。数据可能包含冒犯性或令人不安的内容,仅限于研究用途。数据集格式为JSONL,包括训练/测试分割和详细字段如对话记录、评分等。

This repository provides access to two different kinds of data: 1. Human preference data about helpfulness and harmlessness from Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback. These data are meant to train preference (or reward) models for subsequent RLHF training and are not meant for supervised training of dialogue agents. 2. Human-generated and annotated red teaming dialogues from Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned. These data are meant to understand how crowdworkers red team models and what types of red team attacks are successful or not, and are not meant for fine-tuning or preference modeling. The data contain potentially offensive content and are intended for research purposes. The format is JSONL with splits and fields such as transcripts and ratings.

提供机构:
ZQW329
二维码
社区交流群
二维码
科研交流群
商业服务