DeepSeek-V4-Pro-distill-V2
收藏资源简介:
DeepSeek-V4-Pro-distill-V2是一个用于文本生成任务的蒸馏数据集,包含39,830个通用聊天和指令跟随样本。这些样本通过DeepSeek-V4-Pro API生成响应,并经过GPT-5.5 Thinking结合网络搜索进行事实核查、幻觉检测以及代码示例中的语法和运行时错误修补。数据集采用标准聊天格式,每个样本为一个JSON对象,包含一个messages数组,其中消息角色通常为user或assistant,内容为文本。数据包括单轮对话和多轮对话样本。该数据集覆盖广泛的任务类型,如事实问答、编码、创意写作、解释、翻译、改写、日常对话和多轮助手交互。数据反映了DeepSeek-V4-Pro的风格和知识,可直接用于监督微调(SFT)或作为进一步过滤和混合的基础,但不包含推理内容。数据集基于MIT许可证,可自由使用。
DeepSeek-V4-Pro-distill-V2 is a distilled dataset for text generation tasks, containing 39,830 general chat and instruction-following samples. The responses of these samples are generated using the DeepSeek-V4-Pro API, and subsequently subjected to fact-checking, hallucination detection, and correction of syntax and runtime errors in code examples via GPT-5.5 Thinking integrated with web search. The dataset follows standard chat format, where each sample is a JSON object containing a `messages` array. Messages typically have roles of `user` or `assistant`, with their content being plain text. The dataset includes both single-turn and multi-turn dialogue samples. It covers a wide spectrum of task categories, such as factual question answering, coding, creative writing, explanation, translation, paraphrasing, daily conversations, and multi-turn assistant interactions. The dataset reflects the style and knowledge of DeepSeek-V4-Pro, and can be directly utilized for supervised fine-tuning (SFT) or serve as a foundation for further filtering and mixing, while it does not include reasoning content. The dataset is released under the MIT License and can be freely used.




