Entony12/smoltalk
收藏资源简介:
SmolTalk是一个为大型语言模型(LLMs)的有监督微调(SFT)设计的合成数据集,包含100万个样本,旨在构建SmolLM2-Instruct模型系列。该数据集通过生成新的合成数据来弥补公开SFT数据集的性能差距,覆盖了文本编辑、重写、摘要和推理等多种任务。数据集由多个子集组成,包括新生成的合成数据集(如Smol-Magpie-Ultra、Smol-constraints、Smol-rewrite、Smol-summarize)和现有的公开数据集(如OpenHermes2.5、MetaMathQA、NuminaMath-CoT、Self-Oss-Starcoder2-Instruct、SystemChats2.0、LongAlign、Everyday-conversations、APIGen-Function-Calling、Explore-Instruct-Rewriting),以增强模型在数学、编码、系统提示遵循和长上下文理解等方面的能力。所有新数据集均使用distilabel工具生成,并遵循Apache 2.0许可证。
SmolTalk is a synthetic dataset designed for supervised fine-tuning (SFT) of large language models (LLMs), containing 1 million samples, aimed at building the SmolLM2-Instruct model family. This dataset generates novel synthetic data to compensate for the performance gaps of public SFT datasets, covering a wide range of tasks including text editing, rewriting, summarization and reasoning. The dataset comprises multiple subsets, including newly generated synthetic datasets (e.g., Smol-Magpie-Ultra, Smol-constraints, Smol-rewrite, Smol-summarize) and existing public datasets (e.g., OpenHermes2.5, MetaMathQA, NuminaMath-CoT, Self-Oss-Starcoder2-Instruct, SystemChats2.0, LongAlign, Everyday-conversations, APIGen-Function-Calling, Explore-Instruct-Rewriting), to enhance the model's capabilities in mathematics, coding, system prompt following and long-context understanding. All newly created datasets are generated using the distilabel toolkit and are licensed under Apache 2.0.



