遇见数据集

iamplus/Instruction_Tuning

收藏
Hugging Face2023-05-22 更新2024-03-04 收录
官方服务:

资源简介:

该数据集包含多个子数据集,主要用于指令调优、文章摘要、邮件回复、邮件线程摘要、模型失败案例、身份识别、代码生成、角色扮演、生物学、化学、物理学、数学、金融问答、翻译、自动推理链(COT)等领域。数据集来源包括ChatGPT API、GPT-4 API、以及多个公开数据集。具体数据集包括:IAMAI的种子任务、指令调优数据集、文章摘要数据集、邮件摘要数据集、邮件回复数据集、邮件线程摘要数据集、模型失败案例数据集、身份识别数据集、ChatGPT提示数据集、Stanford Alpaca指令调优数据集、代码生成数据集、ColossalChat指令调优数据集、Laion高质量指令调优数据集、Databricks Dolly人类创建指令调优数据集、GPT-4指令数据集、GPT-4角色扮演数据集、生物学指令数据集、化学指令数据集、物理学指令数据集、数学指令数据集、金融问答指令数据集、翻译指令数据集、自动推理链(COT)指令数据集等。

This dataset contains multiple sub-datasets, primarily targeting applications in instruction tuning, article summarization, email reply generation, email thread summarization, model failure cases, identity recognition, code generation, role-playing, biology, chemistry, physics, mathematics, financial question answering, translation, and automatic chain-of-thought (COT) reasoning. The dataset sources include ChatGPT API, GPT-4 API, and multiple public datasets. Specific included sub-datasets are as follows: IAMAI seed tasks, instruction tuning datasets, article summarization datasets, email summarization datasets, email reply datasets, email thread summarization datasets, model failure case datasets, identity recognition datasets, ChatGPT prompt datasets, Stanford Alpaca instruction tuning dataset, code generation datasets, ColossalChat instruction tuning dataset, Laion high-quality instruction tuning dataset, Databricks Dolly human-created instruction tuning dataset, GPT-4 instruction dataset, GPT-4 role-playing dataset, biology instruction datasets, chemistry instruction datasets, physics instruction datasets, mathematics instruction datasets, financial question answering instruction datasets, translation instruction datasets, and automatic chain-of-thought (COT) reasoning instruction datasets.

提供机构:
iamplus
原始信息汇总

数据集概述

主要数据集

  1. iamai_seed_tasks_v1.csv

    • 内容: IAMAIs seed tasks - Version 1
    • 大小: 879
  2. iamai_v1.csv

    • 内容: Instruction Tuning Dataset collected using seeds from iamai_seed_tasks_v1.csv and ChatGPT API for both prompts and outputs
    • 大小: ~248k
  3. iamai_summarization_v1.csv

    • 内容: Article Summarization dataset (both prompts and outputs) collected using ChatGPT API
    • 大小: ~1.2k
  4. iamai_email_summarization.csv

    • 内容: Email Summarization dataset (both prompts and outputs) collected using ChatGPT API
    • 大小: ~14k
  5. iamai_email_reply_v1.csv

    • 内容: Instruction Tuning Dataset for Email Replying, used ChatGPT API for both prompts and outputs(reply emails)
    • 大小: ~14k
  6. iamai_email_threads.csv

    • 内容: Instruction Tuning Dataset for Email Threads Summarization, used ChatGPT API for both prompts and outputs(thread summaries)
    • 大小: ~17.5k
  7. iamai_failures_v1.csv

    • 内容: Instruction Tuning Dataset collected from failures of model (manojpreveen/gpt-neoxt-20b-v6) and ChatGPT API for outputs
    • 大小: ~10.7k
  8. iamai_identity.csv

    • 内容: Instruction Identity dataset focused on i.am+ organization
    • 模型名称: i.am.ai
    • 组织名称: iam+
    • 大小: ~900

其他相关数据集

  1. chat_gpt_v2.csv

    • 内容: Clean unique prompts collected from external datasets and outputs from ChatGPT API
    • 大小: ~23.8k
  2. stanford_alpaca_it_v3.csv

    • 内容: Instruction Tuning Set with inputs from external set and Outputs from ChatGPT API
    • 大小: ~51.5k
  3. stanford_alpaca_it_v4.csv

    • 内容: Instruction Tuning Set with inputs from external set and Outputs from GPT-4 API
    • 大小: ~51.5k
  4. code_alpaca.csv

    • 内容: Instruction Tuning Set generated Alpaca way for Coding domain with inputs from external set and Outputs from ChatGPT API
    • 大小: ~20k
  5. ColossalChat.csv

    • 内容: Instruction Tuning Set (English) with inputs from external set and Outputs from ChatGPT API
    • 大小: ~52k
  6. unified_chip2.csv

    • 内容: High Quality Instruction Tuning Set by Laion with Python Programming questions split across various programming languages and Outputs from ChatGPT API
    • 大小: ~210k
  7. databricks-dolly.csv

    • 内容: High Quality Human created Instruction Tuning Dataset by Databricks
    • 大小: ~15k
  8. gpt4_instruct.csv

    • 内容: Instruction dataset with outputs from GPT-4
    • 大小: ~18k
  9. gpt4_roleplay.csv

    • 内容: Instruction Roleplay dataset with outputs from GPT-4
    • 大小: ~3k
  10. gpt4_roleplay_v2.csv

    • 内容: Instruction Roleplay Supplemental dataset with outputs from GPT-4
    • 大小: ~7.2k
  11. camel_biology.csv

    • 内容: Instruction dataset on Biology domain with outputs from GPT-4
    • 大小: ~20k
  12. camel_chemistry.csv

    • 内容: Instruction dataset on Chemistry domain with outputs from GPT-4
    • 大小: ~20k
  13. camel_physics.csv

    • 内容: Instruction dataset on Physics domain with outputs from GPT-4
    • 大小: ~20k
  14. camel_math.csv

    • 内容: Instruction dataset on Math domain with outputs from GPT-4
    • 大小: ~50k
  15. FiQA_google.csv

    • 内容: Instruction Tuning dataset on Finance domain with prompts collected from external dataset and outputs from ChatGPT API
    • 大小: ~7k
  16. COIG_translate_en.csv

    • 内容: Instruction Tuning dataset with prompts collected from external dataset and outputs from ChatGPT API
    • 大小: ~66.2k
  17. synthetic_instruct.csv

    • 内容: Instruction Tuning dataset with prompts collected from external dataset and outputs from ChatGPT API
    • 大小: ~33.1k
  18. FLAN_auto_cot.csv

    • 内容: Instruction Tuning dataset (Mainly focused on Math COT) with prompts collected from external dataset and outputs from ChatGPT API
    • 大小: ~8.7k
  19. FLAN_cot_data.csv

    • 内容: Instruction Tuning COT dataset (from FLAN) with prompts collected from external dataset and outputs from ChatGPT API
    • 大小: ~73.4k
  20. LaMini_instruction.csv

    • 内容: Instruction Tuning dataset with prompts from various existing resources of prompts and outputs created using ChatGPT API
    • 大小: ~2.58M
  21. alpaca_evol_instruct_70k.csv

    • 内容: Instruction Tuning dataset - training data of WizardLM
    • 大小: ~70k
搜集汇总
数据集介绍
构建方式
该数据集由iamplus团队精心构建,旨在为指令微调(Instruction Tuning)提供大规模、多领域的高质量训练语料。其构建过程融合了多种策略:首先,基于自研的879条种子任务(iamai_seed_tasks_v1.csv),通过调用ChatGPT API生成约24.8万条指令-响应对,构成了核心数据集iamai_v1。此外,团队从公开数据源如Stanford Alpaca、CodeAlpaca、ColossalChat及LaMini-instruction中筛选并清洗原始提示,再借助ChatGPT或GPT-4 API生成输出,形成了涵盖编程、科学、金融等多领域的子集。针对特定任务,还构建了邮件摘要、回复及对话摘要等专用数据集。所有数据均经过URL剔除、非ASCII字符清理等后处理步骤,以确保数据纯净度。
使用方法
数据集以CSV格式提供,每一条记录包含指令(instruction)与对应输出(output),可直接用于序列到序列模型的监督式微调。用户可通过HuggingFace Datasets库轻松加载,例如使用`load_dataset('iamplus/Instruction_Tuning')`获取全部子集。对于特定任务,如摘要或回复生成,可依据文件名筛选对应子集(如iamai_summarization_v1.csv)。建议在训练前对数据进行分词、批处理等常规预处理,并注意部分子集(如unified_chip2)存在提示重复现象,可根据实际需求进行去重。该数据集兼容主流的深度学习框架,如PyTorch与TensorFlow,适用于对话系统、任务型助手及垂直领域模型的微调实验。
背景与挑战
背景概述
指令微调(Instruction Tuning)作为提升大语言模型遵循人类意图能力的核心技术,近年来备受关注。iamplus/Instruction_Tuning数据集由IAM AI团队于2023年构建,旨在整合多源异构的指令数据,为模型提供涵盖通用对话、代码生成、科学推理、金融问答及邮件处理等领域的多样化训练样本。该数据集通过种子任务扩展、现有公开数据集清洗与重标注、以及基于ChatGPT和GPT-4的合成数据生成等策略,汇聚了超过300万条高质量指令-输出对,为研究指令微调的数据规模、质量与多样性对模型性能的影响提供了重要资源。其发布推动了开源社区对指令数据构建方法的探索,成为连接基础模型与实用化智能助手的关键桥梁。
当前挑战
该数据集面临的核心挑战包括:1) 领域问题层面,如何确保指令数据的多样性与覆盖度以应对模型在长尾任务上的泛化能力不足,例如金融、生物学等专业领域的知识准确性需依赖领域专家验证。2) 构建过程中,依赖大型语言模型(如ChatGPT、GPT-4)生成输出存在潜在偏差与幻觉风险,且不同来源数据(如LaMini的2.58M条样本)存在大量重复提示(约76k条),需高效去重与质量筛选。3) 多语言与跨文化场景下,非英语字符清洗可能导致信息损失,而翻译类指令的处理策略尚需优化。此外,模型失败案例的收集与修正(iamai_failures_v1)虽具创新性,但如何系统化利用这类负样本提升鲁棒性仍是开放问题。
常用场景
经典使用场景
在自然语言处理与大型语言模型飞速发展的当下,指令微调(Instruction Tuning)已成为提升模型遵循人类意图能力的关键技术。IAMAI Instruction Tuning数据集正是为此而生,其经典使用场景聚焦于对基础语言模型进行监督式微调,使其能够精准理解并执行多样化的自然语言指令。该数据集汇集了来自斯坦福Alpaca、GPT-4-LLM、CodeAlpaca、ColossalChat、Dolly等多个知名开源项目的数百万条指令-输出对,覆盖了从通用问答、代码生成到角色扮演、领域知识(如生物、化学、数学)等广泛任务。研究者通常利用该数据集对LLaMA、GPT-NeoX等基座模型进行全参数或LoRA微调,从而赋予模型出色的指令跟随能力,使其在零样本或少样本场景下展现出更强的泛化性与可控性。
解决学术问题
该数据集系统性地解决了学术研究中长期存在的‘指令理解与对齐’难题。传统语言模型虽能生成流畅文本,却常因无法准确捕捉用户意图而产生偏离指令的输出,尤其在多轮对话、复杂推理或专业领域任务中表现欠佳。通过提供海量、高质量且经过人工或强模型(如GPT-4)验证的指令-响应配对,IAMAI Instruction Tuning为研究者构建了统一的基准,用于探究不同微调策略(如数据规模、多样性、输出质量)对模型对齐效果的影响。其意义在于推动了‘指令微调’这一范式的标准化与规模化验证,使得学术界能够系统性地评估模型在遵循指令、抑制幻觉以及跨任务迁移等方面的能力,为后续涌现的InstructGPT、ChatGPT等对齐技术奠定了坚实的实验基础。
实际应用
在实际产业应用中,该数据集是打造高效、可控对话式AI助手的核心基石。企业可基于该数据集微调私有化部署的语言模型,以构建面向客服、教育辅导、代码辅助、金融咨询等垂直场景的智能交互系统。例如,利用其包含的邮件摘要与回复子集(iamai_email_summarization、iamai_email_reply_v1),可快速训练出能够自动提炼邮件要点并生成得体回复的办公自动化工具;借助其代码指令数据(code_alpaca),能开发出理解自然语言描述并生成对应代码的编程助手。此外,通过整合其多领域数据(如金融FiQA、科学Camel系列),可构建跨行业的知识问答平台,显著降低人工服务成本,同时提升用户交互的流畅度与任务完成的准确率。
数据集最近研究
最新研究方向
当前,指令微调(Instruction Tuning)作为大语言模型对齐人类意图的核心技术,正引领着自然语言处理领域的范式革新。iamplus/Instruction_Tuning数据集通过聚合来自ChatGPT、GPT-4等多种前沿模型生成的多样化指令-输出对,覆盖了摘要生成、邮件回复、代码编写、科学推理及金融问答等垂直领域,其规模逾百万条样本,为提升模型在复杂任务中的泛化能力与指令遵循精度提供了关键支撑。该数据集融合了斯坦福Alpaca、ColossalChat及LaMini等开源基准,并引入基于模型失败案例的迭代优化策略,呼应了当前研究对鲁棒性和安全对齐的迫切需求。在学术界与工业界竞相探索高效微调路径的背景下,此类大规模、多源异构的指令数据集已成为推动弱监督学习与少样本推理突破的重要基石,其影响力正延伸至自动化工作流与领域专用智能体的构建中。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务