COIG-CQIA
收藏资源简介:
COIG-CQIA全称为Chinese Open Instruction Generalist - Quality is All You Need,是一个开源的高质量指令微调数据集,由零一万物、中科院深圳先进技术研究院和M-A-P等机构构建。该数据集包含48,375个实例,源自22个不同的数据源,覆盖了从通用知识到STEM领域,再到人文学科的广泛领域。COIG-CQIA以中文互联网获取到的问答及文章作为原始数据,经过深度清洗、重构及人工审核构建而成。该数据集受LIMA: Less Is More for Alignment等研究启发,使用少量高质量的数据即可让大语言模型学习到人类交互行为,因此在数据构建中十分注重数据的来源、质量与多样性。该数据集旨在为中文NLP社区提供高质量且符合人类交互行为的指令微调数据。
COIG-CQIA stands for **Chinese Open Instruction Generalist – Quality is All You Need**, an open-source, high-quality instruction-tuning dataset jointly developed by 01.AI, the Shenzhen Institute of Advanced Technology of the Chinese Academy of Sciences, and M-A-P. The dataset contains 48,375 instances drawn from 22 different sources, covering a broad range of domains from general knowledge to STEM fields and the humanities. COIG-CQIA is constructed from Chinese internet Q&A and articles, which have been deeply cleaned, reconstructed, and manually reviewed. Inspired by research such as *LIMA: Less Is More for Alignment*, which shows that a small amount of high-quality data is sufficient for large language models to learn human interaction behaviors, the dataset places great emphasis on data provenance, quality, and diversity. Its goal is to provide the Chinese NLP community with a high-quality instruction-tuning dataset that aligns with human interaction patterns.




