遇见数据集

0-hero/prompt-perfect

收藏
Hugging Face2024-03-10 更新2024-03-04 收录
官方服务:

资源简介:

该数据集包含了使用GPT-3.5和GPT-4模型对多个流行数据集进行评分的结果。这些数据集包括airoboros-2.1、alpaca-gpt4、dolphin等,评分基于“Self-Alignment with Instruction Backtranslation”论文中的提示。每个数据集都有两个额外的列:score和extracted_score,分别表示模型的响应和提取的分数。评分模型包括gpt-3.5-turbo-16k、gpt-3.5-turbo-1106和gpt-3.5-turbo-0125。评分标准分为5个等级,从1(不完全、模糊、离题、有争议或不符合用户要求)到5(完美的AI助手回答,清晰、逻辑性强、易于理解、引人入胜且富有洞察力)。

This dataset contains the results of scoring multiple popular datasets using GPT-3.5 and GPT-4 models. The datasets include airoboros-2.1, alpaca-gpt4, dolphin, among others. The scoring is based on the prompts from the paper *Self-Alignment with Instruction Backtranslation*. Each dataset features two additional columns: `score` and `extracted_score`, which correspond to the model's generated response and the extracted score, respectively. The scoring models used include gpt-3.5-turbo-16k, gpt-3.5-turbo-1106, and gpt-3.5-turbo-0125. The scoring criteria are categorized into 5 levels, ranging from 1 (incomplete, ambiguous, off-topic, controversial, or failing to meet user requirements) to 5 (a perfect AI assistant response that is clear, logical, easy to comprehend, engaging, and insightful).

提供机构:
0-hero
原始信息汇总

数据集概述

基本信息

  • 语言: 英语
  • 大小: 1M<n<10M
  • 标签: 合成, 蒸馏, GPT-4, GPT-3.5

数据集描述

该数据集包含35个经过评分的大型数据集(超过60亿个令牌),使用GPT-3.5系列模型进行评分。每个数据集包含两个额外列:

  • score: 模型响应,包括CoT(如果提供)
  • extracted_score: 从score列中提取的评分,为整数

评分模型

  • gpt-3.5-turbo-16k
  • gpt-3.5-turbo-1106
  • gpt-3.5-turbo-0125

评分数据集

原始评分提示(来自论文)

  • airoboros-2.1
  • alpaca-gpt4
  • dolphin
  • open-platypus
  • orca_mini_v1
  • SlimOrca-Dedup
  • Synthia-1.3
  • wizard_alpaca_dolly_orca

对话评分提示(修改)

  • Capybara
  • ultrachat

评分分布

数据集 5分 4分 3分 2分 1分 0分
dolphin 80.232373 10.841314 2.217159 3.075088 3.63371 0.000356
open-platypus 76.390115 10.779909 3.093156 3.558533 6.178288 0
Capybara 73.57241 12.851431 3.005123 4.117206 6.435087 0.018743
airoboros-2.1 69.869994 26.695312 1.322096 1.076957 1.035641 0
alpaca-gpt4 65.421891 31.797554 1.301823 0.824937 0.653796 0
wizard_alpaca_dolly_orca 63.898674 32.68317 1.752752 0.894614 0.769829 0.00096
ultrachat 50.213948 40.684169 5.741387 2.880979 0.478934 0.000582
orca_mini_v1 46.351518 49.313846 1.568606 1.898745 0.867284 0
Synthia-v1.3 39.262214 52.335033 2.627859 3.38096 2.392252 0.001683
SlimOrca-Dedup 29.987262 55.132314 7.122872 2.998424 4.759127 0

评分提示

原始评分提示(来自论文)

Below is an instruction from an user and a candidate answer. Evaluate whether or not the answer is a good example of how AI Assistant should respond to the user’s instruction. Please assign a score using the following 5-point scale: 1: It means the answer is incomplete, vague, off-topic, controversial, or not exactly what the user asked for. For example, some content seems missing, numbered list does not start from the beginning, the opening sentence repeats user’s question. Or the response is from another person’s perspective with their personal experience (e.g. taken from blog posts), or looks like an answer from a forum. Or it contains promotional text, navigation text, or other irrelevant information. 2: It means the answer addresses most of the asks from the user. It does not directly address the user’s question. For example, it only provides a high-level methodology instead of the exact solution to user’s question. 3: It means the answer is helpful but not written by an AI Assistant. It addresses all the basic asks from the user. It is complete and self contained with the drawback that the response is not written from an AI assistant’s perspective, but from other people’s perspective. The content looks like an excerpt from a blog post, web page, or web search results. For example, it contains personal experience or opinion, mentions comments section, or share on social media, etc. 4: It means the answer is written from an AI assistant’s perspective with a clear focus of addressing the instruction. It provide a complete, clear, and comprehensive response to user’s question or instruction without missing or irrelevant information. It is well organized, self-contained, and written in a helpful tone. It has minor room for improvement, e.g. more concise and focused. 5: It means it is a perfect answer from an AI Assistant. It has a clear focus on being a helpful AI Assistant, where the response looks like intentionally written to address the user’s question or instruction without any irrelevant sentences. The answer provides high quality content, demonstrating expert knowledge in the area, is very well written, logical, easy-to-follow, engaging and insightful. Please first provide a chain of thought brief reasoning you used to derive the rating score, and then write "Score: <rating>" in the last line.

对话评分提示(修改)

Below are a series of user instructions and corresponding candidate answers in a multi-turn conversation. Evaluate whether or not each answer is a good example of how the AI Assistant should respond to the user’s instructions in the context of an ongoing dialogue. Please assign a score using the following 5-point scale: 1: The answer is incomplete, vague, off-topic, controversial, or fails to build upon previous turns in the conversation. It might ignore context provided earlier, repeat information unnecessarily, or deviate from the conversational flow. Examples include missing content that should logically follow from earlier turns, responses that reset the conversation without acknowledging past interactions, or introducing irrelevant or promotional information. 2: The answer addresses the users concerns but misses key elements of context or nuance from previous turns. It might provide a generally correct direction but fails to leverage the multi-turn nature of the conversation, such as not recalling information provided earlier or not sufficiently building upon it. 3: The answer is helpful and acknowledges the multi-turn context but reads more like a series of standalone responses rather than a cohesive conversation. It covers the basic asks from the user across multiple turns but might lack a seamless integration of conversation history or a sense of ongoing dialogue. 4: The answer is well-tailored to a multi-turn conversation, showing awareness of previous interactions and building upon them effectively. It is clear, comprehensive, and maintains a conversational flow, with only minor room for improvement, such as refining the integration of past and current turns or enhancing conversational fluidity. 5: The answer exemplifies perfect handling of a multi-turn conversation by an AI Assistant. It seamlessly integrates information from previous turns, providing high-quality, context-aware responses that demonstrate expert knowledge and maintain a logical, engaging, and insightful dialogue flow throughout. Please first provide a brief chain of thought reasoning you used to derive the rating score, considering how well the AI Assistant maintains and builds upon the conversational context. Then write "Score: <rating>" in the last line.

搜集汇总
数据集介绍
构建方式
该数据集基于《Self-Alignment with Instruction Backtranslation》论文中提出的评分提示方法构建,旨在对35个流行数据集(总计超过60亿个token)进行质量评估。构建过程中,研究者选用了gpt-3.5-turbo-16k、gpt-3.5-turbo-1106和gpt-3.5-turbo-0125三种模型作为评分工具,对包括airoboros-2.1、alpaca-gpt4、dolphin在内的多个数据集进行打分。每个数据集新增了两个字段:'score'字段存储模型生成的包含思维链的原始回复,'extracted_score'字段则从中提取出整数形式的评分结果。评分提示分为原始版本和针对多轮对话的修改版本,分别适用于单轮指令与多轮对话场景,确保了评估的针对性与全面性。
特点
该数据集的核心特点在于其系统化的质量评估框架。通过引入5分制评分标准,从完整性、相关性、AI助手视角的贴合度等多维度对数据样本进行量化评价,使得数据质量得以直观呈现。评分分布表显示,不同数据集在高质量(5分)和中等质量(4分)区间表现差异显著,如dolphin数据集5分占比高达80.23%,而SlimOrca-Dedup仅为29.99%,这为研究者提供了清晰的比较基准。此外,数据集覆盖了单轮指令和多轮对话两种场景,评分提示经过精心设计,要求模型先输出思维链推理过程再给出分数,增强了评分的可解释性与可靠性。
使用方法
使用该数据集时,用户可直接加载HuggingFace上的'0-hero/prompt-perfect',获取已包含评分结果的35个数据集扩充版本。每个样本的'score'列提供了模型生成的详细评分推理,而'extracted_score'列则便于直接进行数值分析。研究者可根据评分分布筛选高质量子集,例如选择仅包含5分或4分样本的数据进行微调。对于需要自定义评分场景的用户,可参考README中提供的原始评分提示和多轮对话评分提示,将其适配至新的数据集或评估任务中。此外,评分结果还可用于对比不同数据集的质量差异,为数据筛选和模型训练策略提供实证依据。
背景与挑战
背景概述
prompt-perfect数据集由研究团队于2023年基于“Self-Alignment with Instruction Backtranslation”论文的思想构建,旨在系统性地评估和筛选大规模指令微调数据集的质量。该数据集整合了超过35个公开数据集,涵盖超过60亿词元的文本内容,并利用GPT-3.5和GPT-4等先进语言模型对每个样本进行评分,从而为构建高质量、对齐性强的AI助手训练数据提供基准。其核心研究问题在于如何通过自动化评分机制识别出与AI助手行为高度一致的响应,进而推动指令微调领域的数据蒸馏与质量筛选技术发展。该数据集的出现,为后续研究者在数据选择、模型对齐和指令跟随能力提升方面提供了重要参考,显著影响了大规模语言模型训练数据的构建范式。
当前挑战
当前prompt-perfect数据集面临的主要挑战包括:首先,在解决领域问题层面,如何确保评分标准能够准确区分高质量AI助手响应与来自论坛、博客等非AI来源的内容,避免因评分维度单一而误判样本价值。其次,构建过程中,多轮对话数据的评分模型需要兼顾对话上下文连贯性,但现有提示词可能无法充分捕获历史交互中的细微依赖关系,导致评分偏差。此外,不同数据集来源的文本风格和领域分布差异巨大,统一评分模型难以对所有类型样本保持公平性,且评分结果受限于GPT模型自身的主观性和知识边界,存在系统性偏见风险。最后,数据规模庞大(超过10亿样本)使得人工验证成本极高,自动化评分与真实人类判断之间的一致性仍需进一步验证和优化。
常用场景
经典使用场景
在大型语言模型(LLM)的指令微调与对齐优化领域,Prompt-Perfect数据集被广泛用于评估和筛选高质量的训练样本。该数据集基于《Self-Alignment with Instruction Backtranslation》论文中的评分提示,对35个开源指令数据集(如Alpaca-GPT4、Open-Platypus等)进行了系统性打分,为每个样本生成了1至5分的质量评分。研究者通常利用其提供的“score”和“extracted_score”字段,从海量指令数据中遴选出得分较高的子集,用于训练更符合人类偏好、更具安全性和有用性的对话模型。这一经典用法有效缓解了原始数据集中噪声样本对模型性能的负面影响,成为指令微调流程中数据清洗与质量控制的标杆工具。
衍生相关工作
Prompt-Perfect数据集衍生了一系列关于数据质量自动评估与模型对齐的经典工作。例如,后续研究借鉴其评分框架,开发了更精细的多维度评分体系(如安全性、有用性、诚实性),并探索了使用更小模型(如Llama-2)替代GPT-4进行高效评分蒸馏。同时,该数据集催生了“指令回译”方法的改进版本,研究者通过分析评分分布,提出了针对低分样本的迭代修正策略。此外,基于该数据集的评分结果,涌现了如“数据重要性重加权”和“困难样本挖掘”等训练技术,这些工作共同推动了指令微调领域从经验性实践向理论化、系统化的演进。
数据集最近研究
最新研究方向
基于指令反向翻译的自对齐评分方法,当前正被广泛应用于大规模合成数据集的自动化质量评估。该数据集通过引入GPT-3.5与GPT-4系列模型,对超过35个主流指令微调数据集(涵盖超过60亿token)进行一致性评分,并创新性地扩展了对话场景下的多轮交互评估维度。这一研究方向与近期大语言模型对齐技术中强调的“自我改进”与“数据蒸馏”趋势紧密呼应,尤其在后训练阶段数据筛选与质量控制的实际需求中展现出关键价值。通过提供标准化评分框架,该工作为构建更可靠、更符合人类偏好的合成数据生态奠定了方法论基础,对推动开源社区中高质量指令数据的自动化生成与评估具有深远意义。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务