HH-RLHF 和 TL;DR
收藏资源简介:
HH-RLHF和TL;DR是两个用于训练偏好优化技术的偏好数据集。HH-RLHF数据集是通过将人类反馈与LLM的标注相结合,经过精心策划的人类反馈来最大化对齐,而TL;DR数据集则用于总结、合规性和定位等下游任务。数据集的创建是通过粗略的LLM对未标注数据进行初始对齐,然后通过奖励模型和迭代的人类注释来改进对齐。这些数据集的应用领域在于提高大型语言模型与用户偏好的对齐度,减少人类注释的努力,并提高模型在下游任务上的性能。
HH-RLHF and TL;DR are two preference datasets used for training preference optimization techniques. The HH-RLHF dataset combines human feedback with LLM annotations, and is meticulously curated with human feedback to maximize alignment with user preferences. The TL;DR dataset is designed for downstream tasks such as summarization, compliance and localization. These datasets are created by first conducting initial alignment on unlabeled data using coarse LLMs, then refining the alignment via reward models and iterative human annotations. Their applications focus on enhancing the alignment between large language models and user preferences, reducing the effort of human annotation, and improving the model's performance on downstream tasks.




