遇见数据集

iAmBoosted/gpt-oss-20b-reasoning-traces

收藏
Hugging Face2026-05-25 更新2026-05-31 收录
官方服务:

资源简介:

GPT-OSS-20B推理轨迹数据集包含3,333个由openai/gpt-oss-20b生成并经过过滤的推理轨迹,以确保推理过程清洁且终止。该数据集旨在将GPT-OSS的紧密推理风格蒸馏到更小的模型中,并作为iAmBoosted/Qwen3.5-9B-OSS-Distilled模型的训练集。每个记录以聊天消息形式配对一个提示和GPT-OSS-20B的完整推理轨迹及最终答案,适用于监督微调(SFT)。数据集覆盖数学、代码、生物学、化学、物理和谜题等领域,语言为英语。构建过程包括从四个开放数据集中收集提示,使用GPT-OSS-20B生成推理轨迹和答案,并通过过滤步骤去除未清洁终止的轨迹。数据集中大多数记录未验证正确性,仅作为推理风格的演示。数据集采用多许可协议,主要基于Apache-2.0和CC-BY-4.0。

The GPT-OSS-20B Inference Trajectory Dataset contains 3,333 filtered inference trajectories generated by openai/gpt-oss-20b, which are filtered to ensure clean and properly terminated inference processes. This dataset aims to distill the tight inference style of GPT-OSS into smaller models, and serves as the training set for the iAmBoosted/Qwen3.5-9B-OSS-Distilled model. Each record pairs a prompt with the complete inference trajectory and final answer from GPT-OSS-20B in the form of chat messages, which is suitable for Supervised Fine-Tuning (SFT). The dataset covers domains including mathematics, code, biology, chemistry, physics, and puzzles, and is in English. The construction process includes collecting prompts from four open datasets, generating inference trajectories and answers using GPT-OSS-20B, and removing uncleanly terminated trajectories via filtering steps. Most records in the dataset have not been verified for correctness, and only serve as demonstrations of inference styles. The dataset is released under multiple licenses, primarily Apache-2.0 and CC-BY-4.0.

提供机构:
iAmBoosted
二维码
社区交流群
二维码
科研交流群
商业服务