JWei05/DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-all33296-n4
收藏资源简介:
这是一个用于Gemma 3数学蒸馏的教师生成的监督微调(SFT)或蒸馏数据集。数据来源于教师模型JWei05/dapo-gemma3-27b-pt-from-step40-seed43(子文件夹step_000040)和提示集JWei05/DAPO-OpenMathInstruct2-34k的训练分割。数据集包含133,184行数据,基于33,296个唯一提示,每个提示生成4个响应,采样参数为温度=1.0、top_p=1.0、top_k=-1、最大令牌数=20480。列包括messages(聊天消息格式的用户提示和教师助手响应)、teacher_log_probs(每个采样响应令牌的教师对数概率)、teacher_token_ids(生成的响应令牌ID,与teacher_log_probs一一对应)和prompt_idx(源训练分割中的行索引)。预期与rl-distill-scripts/main_distill_offpolicy.py一起使用,用于将前向KL SFT蒸馏到较小的Gemma 3 PT模型中。
This is a teacher-generated supervised fine-tuning (SFT) or distillation dataset for Gemma 3 mathematical distillation. The dataset is sourced from the teacher model JWei05/dapo-gemma3-27b-pt-from-step40-seed43 (subfolder step_000040) and the training split of the prompt set JWei05/DAPO-OpenMathInstruct2-34k. It contains 133,184 rows based on 33,296 unique prompts, with 4 responses generated per prompt using the sampling parameters: temperature=1.0, top_p=1.0, top_k=-1, and max tokens=20480. The columns include messages (user prompts and teacher assistant responses in chat message format), teacher_log_probs (teacher log probabilities for each token in the sampled responses), teacher_token_ids (generated response token IDs that have a one-to-one correspondence with teacher_log_probs), and prompt_idx (the row index in the source training split). It is intended to be used alongside rl-distill-scripts/main_distill_offpolicy.py for conducting forward KL SFT distillation into smaller Gemma 3 PT models.



