glm-4.7-Superior-Reasoning-stage1
收藏资源简介:
glm-4.7-Superior-Reasoning-stage1 是一个基于 Alibaba Superior-Reasoning 风格管道构建的 Stage1 推理蒸馏数据集,采用了更强的教师模型 GLM-4.7 进行高质量推理轨迹生成。该数据集主要关注数学领域,包含 1,192 条记录,采用低温度蒸馏(温度 0.6)方法生成。每条数据为 JSON 格式,包含对话内容、用户问题、助手推理回答、领域标签和元数据(训练阶段、采样温度、教师模型等)。适用于 Stage1 推理监督微调、数学思维链行为调优和教师模型交换消融研究。需要注意的是,这是蒸馏数据而非真实证明数据,推理风格继承教师模型偏好,且需遵守上游数据集的许可和使用限制。
glm-4.7-Superior-Reasoning-stage1 is a Stage1 reasoning distillation dataset built upon the Alibaba Superior-Reasoning-style pipeline, which leverages the more powerful teacher model GLM-4.7 to generate high-quality reasoning trajectories. This dataset primarily focuses on the mathematics domain, includes 1,192 data records, and was generated using low-temperature distillation with a temperature of 0.6. Each entry is formatted in JSON, containing dialogue content, user questions, assistant's reasoning responses, domain tags, and metadata (such as training stage, sampling temperature, teacher model, etc.). It is suitable for Stage1 reasoning supervised fine-tuning, mathematical chain-of-thought behavior tuning, and teacher model swap ablation studies. Notably, this is distillation data rather than real proof data, its reasoning style inherits the preferences of the teacher model, and compliance with the licenses and usage restrictions of the upstream dataset is required.



