BCCard/gemma-4-31b-korean-on-policy-150k
收藏资源简介:
该数据集名为韩语On-Policy问答(Gemma 4)— EAGLE-3训练数据,是一个韩语指令/响应对数据集,其中响应由Gemma 4验证器通过on-policy方式重新生成。最初用于重新训练韩语的EAGLE-3推测器,但也适用于通用的韩语指令调优和蒸馏。数据集包含约150,000行数据,语言为韩语,列包括:instruction(字符串,表示问题/指令)、output(字符串,表示验证器生成的响应)和messages(列表,聊天格式)。数据来源于sh2orc/bccard-maywell-jojo0217-markai-lcw99-kendamarron-microsoft数据集(约171万行韩语问答),仅采样了instruction列(约15万行),原始答案被丢弃。响应由Gemma 4模型(如RedHatAI/gemma-4-26B-A4B-it-FP8-Dynamic或BCCard/gemma-4-31B-it-FP8-Dynamic)生成,生成时关闭思考模式,使用温度1.0、top-p 0.95和top-k 64参数。数据集采用Apache 2.0许可证,提示源数据集和生成响应的Gemma 4模型均为Apache 2.0,响应是Gemma 4生成的合成数据,保留了源数据集的原始归属。
This dataset is named Korean On-Policy QA (Gemma 4) — EAGLE-3 Training Data. It is a Korean instruction-response pair dataset where responses are regenerated via on-policy methods by Gemma 4 validators. Originally developed for retraining the Korean EAGLE-3 speculator, it is also suitable for general Korean instruction tuning and distillation. The dataset contains approximately 150,000 rows of Korean-language data, with three columns: instruction (string, representing questions or instructions), output (string, representing responses generated by validators), and messages (list in chat format). The data is sourced from the sh2orc/bccard-maywell-jojo0217-markai-lcw99-kendamarron-microsoft dataset, which has approximately 1.71 million Korean QA pairs. Only the instruction column was sampled (resulting in ~150,000 rows), and the original answers from the source dataset were discarded. Responses are generated by Gemma 4 models, such as RedHatAI/gemma-4-26B-A4B-it-FP8-Dynamic or BCCard/gemma-4-31B-it-FP8-Dynamic, with thinking mode disabled during generation, using parameters temperature=1.0, top-p=0.95, and top-k=64. The dataset is licensed under Apache 2.0. Both the source prompt dataset and the Gemma 4 models used for response generation are licensed under Apache 2.0. The responses are synthetic data generated by Gemma 4, and the original attribution of the source dataset is retained.




