rampisipati/DeepSeek-V4-Distill-8000x
收藏资源简介:
DeepSeek-V4-Distill-8100x是一个用于推理导向蒸馏的监督微调数据集。问题提示来源于Jackrong/GLM-5.1-Reasoning-1M-Cleaned数据集,答案由教师模型DeepSeek-V4-Flash生成。经过清理过程后,发布的训练集包含7,716个高质量的JSONL示例。清理过程移除了实时问题、身份相关问题、过长问题和其他不合适的提示,以使蒸馏集更加稳定。数据集主要用于推理导向的监督微调、使用DeepSeek-V4-Flash教师输出的蒸馏实验,以及聊天风格和输入/输出SFT管道的格式转换实验。数据格式为JSONL,包含对话式和直接输入/输出字段,如id、conversations、input、output、domain和meta等。数据集存在一些局限性,如可能包含教师模型生成的事实错误或推理伪影。
DeepSeek-V4-Distill-8100x is a supervised fine-tuning dataset for reasoning-oriented distillation. The question prompts come from Jackrong/GLM-5.1-Reasoning-1M-Cleaned, and the answers were generated by the teacher model DeepSeek-V4-Flash. After the cleaning process, the released train split contains 7,716 high-quality JSONL examples. The cleaning process removed real-time questions, identity-related questions, overlong questions, and other unsuitable prompts to make the distillation set more stable. The dataset is primarily intended for reasoning-oriented supervised fine-tuning, distillation experiments using DeepSeek-V4-Flash teacher outputs, and format conversion experiments for chat-style and input/output SFT pipelines. The data format is JSONL, containing both conversation-style and direct input/output fields such as id, conversations, input, output, domain, and meta. The dataset has some limitations, such as potential factual errors or reasoning artifacts inherited from the teacher model.




