m-a-p/COIG-CQIA
收藏资源简介:
COIG-CQIA全称为Chinese Open Instruction Generalist - Quality is All You Need,是一个开源的高质量指令微调数据集,旨在为中文NLP社区提供高质量且符合人类交互行为的指令微调数据。数据集以中文互联网获取到的问答及文章作为原始数据,经过深度清洗、重构及人工审核构建而成。本项目受LIMA: Less Is More for Alignment等研究启发,使用少量高质量的数据即可让大语言模型学习到人类交互行为,因此在数据构建中我们十分注重数据的来源、质量与多样性。数据集包含多个子集,如社交媒体&论坛、通用百科、通用NLP任务、考试&试题、人类价值观、中国传统文化、金融&经管领域、医疗领域和法律领域等。每个子集都有详细的数量、来源和构造方式说明。数据集适用于指令微调,训练模型具备响应指令的能力。
COIG-CQIA, whose full name is Chinese Open Instruction Generalist - Quality is All You Need, is an open-source high-quality instruction tuning dataset. It aims to provide high-quality, human-interaction-aligned instruction tuning data for the Chinese NLP community. The dataset is constructed using raw data consisting of question-answer pairs and articles collected from Chinese internet sources, followed by deep cleaning, restructuring, and manual review. This project is inspired by research including *LIMA: Less Is More for Alignment*, which proves that large language models (LLMs) can acquire human interaction capabilities using only a small volume of high-quality data. Therefore, great emphasis is placed on the source, quality, and diversity of data throughout the dataset construction process. The dataset includes multiple subsets, such as social media & forums, general encyclopedias, general NLP tasks, examinations & test questions, human values, Chinese traditional culture, finance & economics, medical domain, and legal domain. Each subset provides detailed descriptions of its size, source, and construction method. This dataset is intended for instruction tuning, to equip trained models with the capability of responding to user instructions.
数据集概述
数据集名称: COIG-CQIA
全称: Chinese Open Instruction Generalist - Quality is All You Need
目的: 提供高质量的中文指令微调数据集,旨在帮助中文NLP社区训练模型以响应指令。
数据来源: 主要来源于中文互联网的问答及文章。
数据处理: 数据经过深度清洗、重构及人工审核。
语言: 中文
数据集大小: 10K<n<100K
任务类别:
- 问答
- 文本分类
- 文本生成
- 文本到文本生成
数据详情
数据格式
json { "instruction": "示例问题或者指令。", "input": "示例问题或指令的补充。", "output": "对输入的回复。", "task_type": { "major": ["问答"], "minor": ["百科问答"] }, "domain": ["百科", "医疗"], "answer_from": "human", "human_verified": true, "copyright": "作者及版权信息。", }
数据字段
instruction: 指令或问题。input: 问题或指令的补充内容。output: 对应的回答。task_type: 任务类型。domain: 领域分类。answer_from: 回答来源(人类或模型)。human_verified: 是否经过人工验证。copyright: 版权信息。
数据分类及数量
- 社交媒体&论坛:总量13935条
- 知乎:8837条
- 豆瓣:3132条
- 小红书:1508条
- Segmentfault:458条
- 通用百科:总量4571条
- 百科文章:980条
- 中国大百科全书:1706条
- wikiHow中文:1876条
- 通用NLP任务:总量3000条
- COIG-PC-Core:3000条
- 考试&试题:总量2897条
- 高考&中考:2000条
- 研究生入学考试:475条
- 逻辑推理题:422条
- 人类价值观:总量1007条
- 100poison:906条
- COIG-human-value:101条
- 中国传统文化:总量503条
- 中华传统文化试题:232条
- 成语释义:112条
- 古诗词撰写:47条
- 文言文互译:112条
- 金融&经管领域:总量11289条
- MBA百科:10689条
- 金融NLP任务:600条
- 医疗领域:总量8537条
- 医疗百科:8351条
- 医疗文章:186条
- 法律领域:总量2645条
- 法律研究生入学考试:2645条
使用建议
用户应注意数据集的风险、偏差和技术限制。




