遇见数据集

FlagInstruct

收藏
OpenDataLab2026-07-12 更新2024-05-09 收录
官方服务:

资源简介:

我们提出了中文开放教学通才 (COIG) 项目,以维护无害,有用且多样化的中文教学语料库。我们欢迎社区中的所有研究人员为语料库做出贡献并与我们合作。我们只发布COIG的第一个芯片,以帮助中国llm在探索阶段的发展,并呼吁更多的研究人员加入我们的建设COIG。我们介绍了手动验证的翻译通用指令语料库,手动注释的考试指令语料库,人类价值对齐指令语料库,多轮反事实校正聊天语料库和leetcode指令语料库。我们提供这些新的指令语料库,以帮助社区在中文LLMs上进行指令调整。这些指令语料库也是如何有效构建和扩展新中文指令语料库的模板工作流程。

We introduce the Chinese Open Instruction Generalist (COIG) project, which is dedicated to curating a harmless, useful, and diverse Chinese instructional corpus. We welcome all researchers from the community to contribute to this corpus and collaborate with us. We only release the first release of COIG to aid the development of Chinese LLMs during their exploratory phase, and call on more researchers to join us in advancing the COIG project. We present five curated instructional corpora: a manually verified translation-based general instructional corpus, a manually annotated exam instructional corpus, a human value-aligned instructional corpus, a multi-turn counterfactually corrected chat corpus, and a LeetCode instructional corpus. We release these novel instructional corpora to support the community in performing instruction tuning for Chinese LLMs. These corpora also serve as a template workflow for efficiently constructing and scaling new Chinese instructional corpora.

提供机构:
OpenDataLab
创建时间:
2023-10-11
搜集汇总
数据集介绍
FlagInstruct 数据集图片
背景与挑战
背景概述
FlagInstruct是中文开放教学通才(COIG)项目的一部分,旨在提供无害、有用且多样化的中文教学语料库,包含手动验证的翻译指令、考试指令等多种指令数据,以支持中文大型语言模型的指令调整。该数据集由北京智源人工智能研究院于2023年发布,采用Apache 2.0许可证。
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务