InstructionWild v1
收藏资源简介:
The InstructionWild v1 dataset furnishes 52K instructions in both Chinese and English. Constructed using a modelgenerated approach, the dataset involves providing five example prompts to the model, which then generates new instructions along with corresponding responses. The dataset is intended for non-commercial research purposes.
《InstructionWild v1》数据集提供了中英双语共计5.2万条指令。该数据集采用模型生成方式构建,具体流程为向模型输入五条示例提示词,由模型生成全新指令及对应回复。本数据集仅用于非商业性研究用途。
Instruction in the Wild: A User-based Instruction Dataset
数据集概述
- 数据集名称: Instruction in the Wild
- 版本: v1 和 v2
- 数据量:
- v1: 429 条指令
- v2: 超过 110K 条高质量用户指令
- 语言: 英语和中文
- 数据来源: 从 ChatGPT 使用分享中收集的指令
- 数据格式: 与 Alpaca 数据集相同,无输入字段
数据集特点
- 多样性: 数据集中的指令非常多样化,涵盖了生成、开放式问答和头脑风暴等类型。
- 数据收集方法:
- v1: 从 Twitter 上抓取了 700 多条噪声指令,筛选出 429 条高质量指令。
- v2: 未使用自指导生成指令,所有指令均为用户生成。
- 数据标注: v2 版本中对部分指令进行了指令类型和特殊标签的标注。
数据集应用
- 模型训练: Colossal AI 使用该数据集训练了 ColossalChat 模型。
- 模型表现:
- 优点: 在生成、开放式问答和头脑风暴等指令类型上表现较好。
- 局限性:
- 缺乏计数能力、逻辑推理能力、多轮对话和角色扮演能力。
- 在安全性方面存在不足,无法完全遵守 OpenAI 的政策。
数据集对比
- 详细对比: 参见 comparison.md
未来计划
- 待完成: 更大的数据集
作者
引用
bibtex @misc{instructionwild, author = {Jinjie Ni and Fuzhao Xue and Kabir Jain and Mahir Hitesh Shah and Zangwei Zheng and Yang You }, title = {Instruction in the Wild: A User-based Instruction Dataset}, year = {2023}, publisher = {GitHub}, journal = {GitHub repository}, howpublished = {url{https://github.com/XueFuzhao/InstructionWild}}, }




