遇见数据集

COIG-CQIA

收藏
OpenCSG2023-12-28 更新2026-01-19 收录
官方服务:

资源简介:

COIG-CQIA是一个开源的高质量中文指令微调数据集,致力于为中文NLP社区提供符合人类交互行为的数据。它主要包含从中文互联网上获取的问答及文章,经过清洗、重构和人工审核。该数据集包含社交媒体、论坛、通用百科、通用NLP任务、考试试题、人类价值观、中国传统文化、金融经管和医疗等多个领域的数据,总量在1万到10万之间。数据格式为JSON,包含指令、输入、输出、任务类型、领域、回答来源、人工核验和版权信息等字段,适用于指令微调任务,旨在训练模型具备响应指令的能力。

COIG-CQIA is an open-source, high-quality Chinese instruction-tuning dataset dedicated to providing data that aligns with human interaction behaviors for the Chinese NLP community. It primarily comprises question-answering pairs and articles collected from Chinese internet platforms, which have undergone cleaning, restructuring, and manual review. This dataset covers multiple domains including social media, forums, general encyclopedias, general NLP tasks, examination questions, human values, traditional Chinese culture, financial economics and management, and medical care, with a total scale ranging from 10,000 to 100,000 samples. The data is stored in JSON format, containing fields such as instruction, input, output, task type, domain, answer source, manual verification, and copyright information. It is tailored for instruction tuning tasks, with the goal of training models to possess the capability to respond to user instructions properly.

提供机构:
AIWizards
创建时间:
2024-04-18
搜集汇总
数据集介绍
COIG-CQIA 数据集图片
以上内容由遇见数据集搜集并总结生成
二维码
社区交流群
二维码
科研交流群
商业服务