anshy/Superior-Reasoning-SFT-gpt-oss-120b-random-shuffled
收藏资源简介:
Superior-Reasoning-SFT-gpt-oss-120b数据集是一个高质量的开源集合,包含43.5万个样本,旨在民主化高性能长链思维推理模型的训练。该数据集采用分布对齐序列蒸馏管道构建,解决了当前推理蒸馏中的关键限制,特别是分布不匹配和暴露偏差问题,使得小型密集模型能够实现与规模相当的开源模型中的最先进性能,并显著超越更大模型。数据集覆盖数学、代码生成、科学推理和指令遵循等多个领域,所有数据均从gpt-oss-120b模型的高推理模式中蒸馏而来,确保高质量的推理轨迹。数据集分为两个阶段:阶段1(低温度训练,约10.5万样本)用于稳定性训练,阶段2(高温度训练,约33万样本)用于多样性训练。数据来源于多个公开数据集,如nvidia/AceReason-1.1-SFT等,并采用严格的质控措施。数据集可用于复现DASD-4B-Thinking等模型的性能,在多个基准测试中表现优异。
The Superior-Reasoning-SFT-gpt-oss-120b dataset is a high-quality, open-source collection containing 435K samples designed to democratize the training of high-performance Long Chain-of-Thought models. It is constructed using a principled Distribution-Aligned Sequence Distillation pipeline, addressing key limitations in current reasoning distillation—specifically distributional mismatch and exposure bias—enabling small dense models to achieve State-of-the-Art performance among open-source models of comparable scale and outperform significantly larger models. The dataset spans diverse domains including Mathematics, Code Generation, Scientific Reasoning, and Instruction Following, with all data distilled from gpt-oss-120b using its high-reasoning mode to ensure high-quality reasoning traces. It is divided into two stages: Stage 1 (low-temperature training, ~105K samples) for stability and Stage 2 (high-temperature training, ~330K samples) for diversity. Data sources include publicly available datasets such as nvidia/AceReason-1.1-SFT, and rigorous quality control measures are applied. The dataset can be used to reproduce the performance of models like DASD-4B-Thinking and demonstrates excellent results on multiple benchmarks.



