openai/summarize_from_feedback
收藏资源简介:
在《Learning to Summarize from Human Feedback》论文中,研究人员从人类反馈中训练了一个奖励模型,该模型随后用于训练摘要模型以符合人类偏好。此数据集是为此奖励模型发布的人类反馈数据。数据集分为两部分:comparisons和axis。在comparisons部分,人类注释者被要求从两个摘要中选择最佳摘要。在axis部分,人类注释者对摘要的质量进行了评分。comparisons部分仅包含训练和验证分割,而axis部分仅包含测试和验证分割。用于训练奖励模型的摘要来自TL;DR数据集,额外的验证和测试数据来自TL;DR数据集、CNN文章和Daily Mail文章。
In the paper *Learning to Summarize from Human Feedback*, researchers trained a reward model using human feedback, which was subsequently used to fine-tune summarization models to align with human preferences. This dataset is the human feedback data released for training this reward model. The dataset is divided into two subsets: comparisons and axis. In the comparisons subset, human annotators are required to select the best summary from two given summaries. In the axis subset, human annotators rate the quality of summaries. The comparisons subset only includes training and validation splits, while the axis subset only contains test and validation splits. The summaries used for training the reward model are sourced from the TL;DR dataset, and additional validation and test data are derived from the TL;DR dataset, CNN articles, and Daily Mail articles.
数据集概述
数据集名称
- pretty_name: Summarize from Feedback
数据集描述
- 来源与目的: 该数据集源自论文《Learning to Summarize from Human Feedback》,用于训练奖励模型,进而训练出符合人类偏好的摘要模型。
- 数据集组成:
- comparisons: 人类标注者从两个摘要中选择最佳的一个。
- axis: 人类标注者对摘要的质量进行Likert量表评分。
- 数据集分割:
comparisons部分包含训练集和验证集。axis部分包含测试集和验证集。
- 数据来源: 训练奖励模型的摘要数据来自TL;DR数据集,额外的验证和测试数据来自TL;DR数据集、CNN文章和Daily Mail文章。
引用信息
- 论文: Learning to Summarize from Human Feedback
- 作者: Nisan Stiennon, Long Ouyang, Jeff Wu, Daniel M. Ziegler, Ryan Lowe, Chelsea Voss, Alec Radford, Dario Amodei, Paul Christiano
- 发表年份: 2020
- 会议: NeurIPS




