遇见数据集

JWei05/DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4

收藏
Hugging Face2026-05-10 更新2026-05-31 收录
官方服务:

资源简介:

DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4是一个用于Gemma 3数学蒸馏的教师生成监督微调或蒸馏数据集。数据来源于教师模型JWei05/dapo-gemma3-27b-pt-from-step40-seed43(子文件夹step_000040)和提示集JWei05/DAPO-OpenMathInstruct2-34k(训练分割)。数据集包含128,000行数据,基于32,000个唯一提示,每个提示生成4个响应,采样时使用温度参数1.0、top_p=1.0、top_k=-1和最大token数20480。主要列包括:messages(以聊天消息格式存储的用户提示和教师助理响应)、teacher_log_probs(教师对每个采样响应token的对数概率)、teacher_token_ids(生成的响应token ID,与teacher_log_probs一一对应)和prompt_idx(源训练分割中的行索引)。预期用途是与rl-distill-scripts/main_distill_offpolicy.py脚本结合,用于将前向KL监督微调蒸馏到较小的Gemma 3 PT模型中。

DAPO-Gemma3-27B-PT-RL-step40-seed43-SFT-Data-32k-n4 is a teacher-generated SFT/distillation dataset for Gemma 3 math distillation. The data is sourced from the teacher model JWei05/dapo-gemma3-27b-pt-from-step40-seed43 (subfolder step_000040) and prompts from JWei05/DAPO-OpenMathInstruct2-34k (train split). It contains 128,000 rows, based on 32,000 unique prompts with 4 responses per prompt, sampled with temperature=1.0, top_p=1.0, top_k=-1, and max_tokens=20480. Key columns include: messages (user prompt and teacher assistant response in chat-message format), teacher_log_probs (teacher log-probability of each sampled response token), teacher_token_ids (generated response token IDs, aligned 1:1 with teacher_log_probs), and prompt_idx (row index in the source train split). It is intended for use with rl-distill-scripts/main_distill_offpolicy.py for forward-KL SFT distillation into smaller Gemma 3 PT models.

提供机构:
JWei05
二维码
社区交流群
二维码
科研交流群
商业服务