berkeley-nest/Nectar
收藏资源简介:
Nectar是第一个通过GPT-4生成的高质量7-wise比较数据集,包含多样化的聊天提示、高质量和多样化的响应以及准确的排名标签。数据集的提示来源于多个数据集,包括lmsys-chat-1M、ShareGPT、Antropic/hh-rlhf、UltraFeedback、Evol-Instruct和Flan。每个提示的7个响应主要来自GPT-4、GPT-3.5-turbo、GPT-3.5-turbo-instruct、LLama-2-7B-chat和Mistral-7B-Instruct等模型。数据集的响应通过GPT-4进行排序,总共包含3.8M对比较。Nectar用于训练奖励模型Starling-RM-7B-alpha,该模型推动了Starling-LM-7B-alpha在MT-Bench上的得分达到8.09,这是目前任何7B模型的最高得分。
Nectar is the first high-quality 7-wise comparison dataset generated via GPT-4, featuring diverse chat prompts, high-quality and varied responses, as well as accurate ranking labels. The prompts of this dataset are sourced from multiple existing datasets, including lmsys-chat-1M, ShareGPT, Antropic/hh-rlhf, UltraFeedback, Evol-Instruct and Flan. For each prompt, its 7 responses are primarily generated by models such as GPT-4, GPT-3.5-turbo, GPT-3.5-turbo-instruct, LLama-2-7B-chat and Mistral-7B-Instruct. All responses in the dataset are ranked using GPT-4, with a total of 3.8 million comparison pairs. Nectar is utilized to train the reward model Starling-RM-7B-alpha, which enabled Starling-LM-7B-alpha to achieve a score of 8.09 on MT-Bench, the highest score attained by any 7B model to date.
数据集概述
基本信息
- 名称: Nectar
- 开发团队: Banghua Zhu, Evan Frick, Tianhao Wu, Hanlin Zhu, Jiantao Jiao
- 许可: Apache-2.0,条件是不用于与OpenAI竞争
- 语言: 英语
- 大小分类: 100K<n<1M
- 配置: 默认配置,数据文件位于
data/rlaif.parquet - 标签: RLHF, RLAIF, reward model
数据集内容
- 类型: 7-wise比较数据集
- 生成方式: 通过GPT-4基于排名的生成
- 内容: 包含多样化的聊天提示、高质量和多样化的响应以及精确的排名标签
- 提示来源: 混合了多个数据集,包括lmsys-chat-1M, ShareGPT, Antropic/hh-rlhf, UltraFeedback, Evol-Instruct, Flan
- 响应来源: 主要来自GPT-4, GPT-3.5-turbo, GPT-3.5-turbo-instruct, Llama-2-7B-chat, Mistral-7B-Instruct等模型
- 排名: 每个提示的7个响应由GPT-4进行排名,总计3.8M对比较
数据集结构
json { "prompt": str, "answers": [ { "answer": str, "model": str, "rank": int }, ... ], "turns": int, "num_response": int, "source": list[str], "good_natured": bool }
数据收集过程
- 提示收集: 从多个数据集生成并筛选提示,确保每个提示有多个答案
- 响应收集: 从多个模型中提取响应,确保每个提示有7个响应
- 排名收集: 使用GPT-4根据帮助性和无害性标准对响应进行排名
注意事项
- 数据集包含可能不安全、冒犯性或令人不安的内容,仅供训练更安全的模型使用。




