SYNTH-Swallow-Math-Code-Mix
收藏资源简介:
该数据集是一个高质量的全合成数据集,由四个来源的数据完全混洗而成:SYNTH(约63.5%)、SwallowCode-v2(约15.5%)、SwallowMath-v2-textbook(约10.5%)和SwallowMath-v2-qa(约10.0%)。创建此数据集的动机是为了提供一个方便、预混洗的高质量合成/增强数据集集合,用于2026年1月前的小型语言模型预训练实验。SYNTH数据集在训练中占主导地位,但辅以tokyotech提供的优秀数学和代码数据集,这些数据集经过重写,具有一致且可预测的格式。通过这种组合,模型可以获得SYNTH设计上缺乏的通用知识(SYNTH)、基础数学理解(Swallow Math V2)和Python代码技能(Swallow Code V2)。数据集采用单列格式(`text`),以保持训练的简洁性。所有荣誉归功于原始作者PleIAs和tokyotech。
This is a high-quality fully synthetic dataset completely shuffled from four sources: SYNTH (approximately 63.5%), SwallowCode-v2 (approximately 15.5%), SwallowMath-v2-textbook (approximately 10.5%), and SwallowMath-v2-qa (approximately 10.0%). The motivation for creating this dataset is to provide a convenient, pre-shuffled high-quality synthetic/augmented dataset collection for small language model pre-training experiments prior to January 2026. The SYNTH dataset dominates the training, but is supplemented by excellent mathematical and code datasets provided by tokyotech. These datasets have been rewritten to feature a consistent and predictable format. Through this combination, the model can acquire general knowledge (lacking in the original SYNTH dataset), basic mathematical understanding from Swallow Math V2, and Python coding skills from Swallow Code V2. The dataset adopts a single-column format (`text`) to maintain training simplicity. All credits go to the original authors PleIAs and tokyotech.
数据集概述
数据集名称
Mixed dataset: SYNTH + SwallowMath-v2 + SwallowCode-v2
数据集来源与构成
本数据集是一个高质量、全合成的混合数据集,由以下四个来源的数据完全随机打乱混合而成:
- SYNTH:占比约 63.5%
- SwallowCode-v2:占比约 15.5%
- SwallowMath-v2-textbook:占比约 10.5%
- SwallowMath-v2-qa:占比约 10.0%
原始数据来源详情
PleIAs/SYNTH:格式为text = query + synthetic_reasoning + synthetic_answertokyotech-llm/swallow-math-v2:包含子集swallow-math-v2-qa和swallow-math-v2-textbooktokyotech-llm/swallow-code-v2:包含子集stage5-auto-format
数据集特征
- 数据列:仅包含单列
text,以保持训练过程极其简单。 - 语言:英语
- 任务类别:文本生成
- 许可协议:其他
创建动机
提供此数据集的目的是为了满足截至2026年1月,在小语言模型预训练实验中,需要一个方便、预先打乱混合的、最高质量的合成/增强数据集的合并版本。SYNTH 数据将在此训练中占主导地位,并辅以 tokyotech 提供的优秀的数学和代码数据集。这些数据集由格式一致且可预测的改写合成样本组成,是对 SYNTH 的很好补充。通过这种组合,模型应能获得通用知识(来自 SYNTH)、基础数学理解(来自 Swallow Math V2)以及 Python 代码技能(来自 Swallow Code V2),而这些是当前形式的 SYNTH 在设计上所缺乏的。
致谢
所有荣誉归于原始作者:PleIAs 和 tokyotech。




