toroe/13m_17t
收藏资源简介:
13m_17t SFT Blend是一个大规模监督微调(SFT)混合数据集,包含约13.6百万个指令遵循示例,涵盖数学、科学、代码、通用聊天、指令遵循、工具调用和安全等多个领域。数据集于2026年5月19日通过去重感知的UID交集采样从Nemotron系列数据集中构建,确保无跨源重复。数据格式为JSONL,采用OpenAI风格的聊天消息格式(包含`messages`、`source`、`dataset_name`、`ds_uid`字段)。主要语言为英语,德语约占17.6%。数据集中混合了推理开启(带有`<think>...</think>`链式推理痕迹)和推理关闭的示例。预期用途是通过Megatron-LM训练框架进行大语言模型的监督微调,以提升在推理密集型领域(如数学、科学、代码)的指令遵循能力,同时支持多语言(德语)覆盖和代理技能(如工具调用、终端代理、软件工程)。局限性包括德语示例为机器翻译且质量未大规模验证、安全覆盖极少、德语翻译切片内未显式去重,以及推理痕迹质量因领域而异。
The 13m_17t SFT Blend is a large-scale Supervised Fine-Tuning (SFT) blend dataset containing approximately 13.6 million instruction-following examples spanning multiple domains including mathematics, science, code, general chat, instruction following, tool calling, and safety. This dataset was constructed from the Nemotron-series datasets on May 19, 2026 via deduplication-aware UID intersection sampling, ensuring no cross-source duplicates. The data is stored in JSONL format, adhering to the OpenAI-style chat message schema with fields including `messages`, `source`, `dataset_name`, and `ds_uid`. The primary language of the dataset is English, with German accounting for approximately 17.6% of the total content. The dataset includes both reasoning-enabled examples (marked with `<think>...</think>` chain-of-thought traces) and reasoning-disabled examples. Its intended use is for supervised fine-tuning of large language models (LLMs) via the Megatron-LM training framework, to enhance instruction-following capabilities in reasoning-intensive domains such as mathematics, science, and code, while also supporting multilingual (German) coverage and agent skills including tool calling, terminal agents, and software engineering. Limitations of the dataset are as follows: German examples are machine-translated without large-scale quality validation, there is minimal safety coverage, no explicit deduplication is performed within German translation slices, and the quality of chain-of-thought traces varies across different domains.



