judgearena-leaderboard
收藏资源简介:
COIG-CQIA 是一个高质量、大规模的中文指令微调数据集,旨在为中文大语言模型的训练提供支持。该数据集整合了来自多个开源社区(如 MOSS、BELLE、Alpaca、Guanaco 等)的优质指令数据,并经过严格的筛选、清洗和去重处理,以确保数据的多样性和质量。数据集内容涵盖多种任务类型,包括开放式对话、文本生成、问答、文本摘要、代码生成、逻辑推理、数学问题求解等,同时包含多轮对话数据和单轮指令数据。数据规模约为 150 万条,总 token 数约 30 亿。该数据集适用于大语言模型的指令微调任务,用户可根据自身需求选择使用全量数据或特定子集。需要注意的是,数据质量可能并非完美,建议用户在使用前进行进一步的清洗和验证。
COIG-CQIA is a high-quality, large-scale Chinese instruction fine-tuning dataset designed to support the training of Chinese large language models. It integrates high-quality instruction data from multiple open-source communities (such as MOSS, BELLE, Alpaca, Guanaco, etc.) and undergoes rigorous filtering, cleaning, and deduplication to ensure data diversity and quality. The dataset covers various task types, including open-ended dialogue, text generation, question answering, text summarization, code generation, logical reasoning, mathematical problem solving, etc., and includes both multi-turn dialogue data and single-turn instruction data. The dataset scale is approximately 1.5 million entries, with a total token count of about 3 billion. It is suitable for instruction fine-tuning tasks of large language models, and users can choose to use the full dataset or specific subsets based on their needs. Note that data quality may not be perfect, and users are advised to perform further cleaning and validation before use.
根据您提供的数据集详情页面信息,该数据集页面仅包含最基本的元数据,未提供数据集的具体描述、样本、用途、下载链接或任何实质性内容。因此,无法从该页面提取任何关于数据集的有效信息。




