Bespoke-Stratos-17k
收藏资源简介:
Bespoke-Stratos-17k是一个推理数据集,包含问题、推理轨迹和答案。该数据集是通过复制和改进Berkeley Sky-T1数据管道,并使用DeepSeek-R1的SFT蒸馏数据创建的。数据集用于训练两个模型:Bespoke-Stratos-32B和Bespoke-Stratos-7B。Bespoke-Stratos-32B是基于Qwen-2.5-32B-Instruct微调的32B推理模型,而Bespoke-Stratos-7B是基于Qwen-2.5-7B-Instruct微调的7B推理模型。数据集的生成过程使用了Bespoke Curator工具,并在1.5小时内完成,成本为800美元。与Sky-T1相比,Bespoke-Stratos-17k使用了DeepSeek-R1作为教师推理模型,并且没有重新格式化DeepSeek-R1的推理轨迹。此外,使用了gpt-4o-mini来过滤错误的数学解决方案,从而提高了正确解决方案的保留率。
Bespoke-Stratos-17k is a reasoning dataset containing questions, reasoning traces and answers. This dataset was created by replicating and improving the Berkeley Sky-T1 data pipeline, using the SFT distilled data of DeepSeek-R1. It is used to train two models: Bespoke-Stratos-32B and Bespoke-Stratos-7B. Bespoke-Stratos-32B is a 32-billion-parameter reasoning model fine-tuned based on Qwen-2.5-32B-Instruct, while Bespoke-Stratos-7B is a 7-billion-parameter reasoning model fine-tuned based on Qwen-2.5-7B-Instruct. The dataset generation process used the Bespoke Curator tool, completed within 1.5 hours with a cost of $800. Compared with Sky-T1, Bespoke-Stratos-17k uses DeepSeek-R1 as the teacher reasoning model without reformatting its reasoning traces. Additionally, gpt-4o-mini was used to filter incorrect mathematical solutions, thereby improving the retention rate of correct solutions.




