smoltalk-chinese
收藏资源简介:
smoltalk-chinese 是一个参考 SmolTalk 数据集构建的中文微调数据集,旨在为大型语言模型(LLM)的训练提供高质量的合成数据支持。该数据集全部由合成数据组成,涵盖超过70万条数据,专门设计用于提升中文大型语言模型在多种任务上的表现,增强模型的多功能性和适应性。数据集由多个部分组成,包括参考magpie-ultra的任务类型、参考smoltalk的其它任务类型、模拟日常生活中的对话风格以及来自Math23K中文版的数学题数据。数据集的生成过程严格遵循高标准,确保数据的质量和多样性。实验验证表明,基于smoltalk-chinese微调的模型在多个指标上表现出显著优势。
smoltalk-chinese is a Chinese fine-tuning dataset developed by referencing the SmolTalk dataset, which aims to provide high-quality synthetic data support for the training of large language models (LLMs). This dataset is entirely composed of synthetic data, containing over 700,000 entries, and is specially designed to enhance the performance of Chinese large language models across diverse tasks, as well as improve their versatility and adaptability. The dataset consists of multiple components, including task types adapted from magpie-ultra, other task types derived from SmolTalk, conversational styles simulating daily life, and mathematical problem data from the Chinese version of Math23K. The dataset generation process strictly adheres to high-quality standards to ensure both the quality and diversity of the data. Experimental validations have shown that models fine-tuned using smoltalk-chinese exhibit significant advantages across multiple evaluation metrics.
Chinese SmolTalk 数据集概述
数据集基本信息
- 语言: 中文 (zh)
- 任务类别: 文本生成 (text-generation)
- 许可证: Apache-2.0
- 数据规模: 10B < n < 100B
数据集描述
smoltalk-chinese 是一个参考 SmolTalk 数据集构建的中文微调数据集,旨在为大型语言模型(LLM)的训练提供高质量的合成数据支持。该数据集全部由合成数据组成,涵盖超过70万条数据,专门设计用于提升中文大型语言模型在多种任务上的表现,增强模型的多功能性和适应性。
数据集组成
-
Magpie-Ultra 参考任务
- 使用 Magpie 合成的三轮对话数据,任务包括:
- 信息检索 (Information-seeking)
- 推理 (Reasoning)
- 规划 (Planning)
- 编辑 (Editing)
- 编程 (Coding)
- 数学 (Math)
- 角色扮演 (Role-playing)
- 数据分析 (Data-analysis)
- 创意写作 (Creative-writing)
- 寻求建议 (Advice-seeking)
- 头脑风暴 (Brainstorming)
- 使用 Magpie 合成的三轮对话数据,任务包括:
-
SmolTalk 参考任务
- 使用 Magpie 合成的一轮对话任务,任务包括:
- 格式约束 (Format-constrain)
- 重写 (Rewrite)
- 总结 (Summary)
- 安全 (Safe)
- 翻译 (Translate)
- 文档问答 (Doc)
- 使用 Magpie 合成的一轮对话任务,任务包括:
-
模拟日常对话
- 生成五轮对话数据,模拟日常生活中的对话风格。
-
数学问题
- 来自 Math23K 中文版的数学题数据,答案包含详细推理步骤,由 deepseek-v2.5 生成。
数据集生成方法
- 数据生成: 使用 Magpie 合成原始数据,生成模型包括 deepseek-v2.5 和 qwen2.5-72b-instruct,结合 Distilabel 库确保生成内容的丰富性和多样性。
- 数据筛选: 利用 qwen2-7b-instruct 模型对对话数据的第一条指令进行清晰度和流畅度评分,仅保留评分在2分及以上的数据。
- 去重处理: 使用 gte-large-zh 模型对对话数据的第一条指令进行编码,根据嵌入相似度进行去重处理,确保数据的独特性和多样性。
实验验证
- 基础模型: 使用 opencsg/csg-wukong-ablation-chinese-fineweb-edu(在 chinese-fineweb-edu 上预训练的2B模型)作为基础模型。
- 微调过程: 在 smoltalk-chinese、Magpie-Qwen2-Pro-200K-Chinese 和 infinity-instruct 数据集上进行微调,训练设置为:
- Epochs: 2
- Learning Rate: 3e-4
- Scheduler: Cosine decay
- Global Batch Size: 32
- 评估结果: 在 Alignbench 上评估模型的中文对话能力,结果表明,基于 smoltalk-chinese 微调的模型在多个指标上表现出显著优势。
许可协议
使用 Chinese SmolTalk 数据集需要遵循 OpenCSG 社区许可证。该数据集支持商业用途,但需发送邮件至 lorraineg@opencsg.com 并获得许可。




