JWei05/DAPO-Gemma3-4B-PT-DAPO-17.4k
收藏资源简介:
该数据集是为DAPO-17.4k生成的教师数据,使用google/gemma-3-4b-pt模型。训练分割包含262,368行,每训练问题有16个响应;验证分割包含1,000行,每验证问题有1个响应。采样采用默认教师生成采样参数,提示使用Gemma 3 IT聊天模板与盒装答案指令。每行数据包括:OpenAI风格聊天消息(messages)、完整聊天标记化的提示和响应标记ID(input_ids)、与input_ids对齐的标记级掩码(response_mask,响应标记标记为1)、仅响应标记ID(teacher_token_ids)、与teacher_token_ids对齐的仅响应标记对数概率(teacher_log_probs)、以及来自源分割的原始提示索引(prompt_idx)。
Teacher generations for DAPO-17.4k using google/gemma-3-4b-pt. Train split: 16 responses per training question, 262,368 rows total. Validation split: 1 response per validation question, 1,000 rows total. Sampling: default teacher generation sampling parameters. Prompting: Gemma 3 IT chat template with the boxed-answer instruction. Each row includes: OpenAI-style chat messages (messages), full chat-tokenized prompt plus response token ids (input_ids), token-level mask aligned to input_ids with response tokens marked as 1 (response_mask), response-only token ids (teacher_token_ids), response-only token log probabilities aligned to teacher_token_ids (teacher_log_probs), and original prompt index from the source split (prompt_idx).



