BAAI/COIG
收藏资源简介:
Chinese Open Instruction Generalist (COIG)项目旨在构建一个无害、有帮助且多样化的中文指令语料库。项目包含多个子数据集,如手动验证的翻译通用指令语料库、手动注释的考试指令语料库、人类价值观对齐指令语料库、多轮反事实修正聊天语料库以及LeetCode指令语料库。这些语料库旨在帮助中文大型语言模型(LLMs)的指令调优,并为构建新的中文指令语料库提供模板。
The Chinese Open Instruction Generalist (COIG) project aims to construct a harmless, helpful and diverse Chinese instruction corpus. The project includes multiple sub-datasets, such as manually verified general translation instruction corpus, manually annotated exam instruction corpus, human values aligned instruction corpus, multi-turn counterfactual revision chat corpus, and LeetCode instruction corpus. These corpora are designed to facilitate instruction tuning for Chinese large language models (LLMs) and provide templates for building new Chinese instruction corpora.
数据集概述
COIG(Chinese Open Instruction Generalist) 项目旨在维护一个无害、有益且多样化的中文指令语料库集合。该项目欢迎社区研究人员的贡献与合作,并已发布首批数据以支持中文大型语言模型的发展。
数据集内容
-
Translated Instructions
包含66,858条指令,源自Super-NaturalInstructions、Self-Instruct和Unnatural Instructions,通过自动翻译、手动验证和手动校正三阶段流程确保质量。 -
Exam Instructions
包含63,532条指令,主要来自中国的高考、中考和公务员考试,涵盖多种题型和详细解析,适用于思维链(CoT)语料库。 -
Human Value Alignment Instructions
包含34,471条指令,分为两类:共享人类价值观和特定地区或国家的价值观。 -
Counterfactural Correction Multi-round Chat
包含13,653条指令,基于CN-DBpedia知识图谱,旨在解决大型语言模型中的幻觉和事实不一致问题。 -
Leetcode Instructions
包含11,737条指令,源自Leetcode编程问题,旨在增强语言模型在代码相关任务上的能力。
数据集下载
建议直接从Hugging Face下载所需数据文件,而非使用HF load_datasets。
许可证
COIG数据集由BAAI发布,遵循Apache 2.0许可证。部分内容可能采用其他许可,如MIT许可证。




