mediumish-small-agent-sft-preview-v0.1
收藏资源简介:
mediumish-small-agent-sft-preview-v0.1 是一个用于中等/小型模型的程序化生成的ChatML监督微调(SFT)数据集预览版。该数据集专门设计用于训练模型在多个关键领域的表现,包括智能体工具使用、抗幻觉习惯培养、有根据的拒绝能力、多步推理以及相关的认知行为。数据集包含6,976个样本,替代了先前5,000行的预览版本。数据在20个不同的应用领域分层分布,确保覆盖多样性,这些领域包括:可靠性与工具使用、圣经研究、隐藏假设推理、高级数学、基于文档规范的代码修复、规则手册与政策模拟、ARC风格网格谜题、状态机、逻辑网格、引用解析、反事实推理等。为防止单一领域主导训练,数据集的字符分布受到刻意限制,例如代码相关内容的字符占比约为28%,而ARC网格交互领域的字符占比低于5%。数据集包含两个主要列:chatml列存储完整的多轮对话训练示例,涵盖系统消息、用户输入、助手回复以及工具调用等轮次(如适用);difficulty列标识每个样本的难度等级,分为easy(简单)、medium(中等)、hard(困难)和adversarial(对抗性)四个级别。该版本明确标注为预览版,可能存在问题,并非最终的生产发布版本。
mediumish-small-agent-sft-preview-v0.1 is a programmatically generated ChatML supervised fine-tuning (SFT) dataset preview for medium/small models. It is specifically designed to train models in multiple key areas, including agent tool usage, anti-hallucination habit cultivation, grounded refusal capabilities, multi-step reasoning, and related cognitive behaviors. The dataset contains 6,976 samples, replacing the previous 5,000-row preview version. Data is stratified across 20 different application domains to ensure diversity, covering areas such as: reliability and tool usage, biblical studies, hidden assumption reasoning, advanced mathematics, code fixes based on documentation specifications, rulebook and policy simulation, ARC-style grid puzzles, state machines, logic grids, citation resolution, counterfactual reasoning, and more. To prevent any single domain from dominating training, the character distribution is deliberately limited, with code-related content comprising about 28% of characters, while ARC grid interaction domains account for less than 5%. The dataset includes two main columns: the chatml column stores complete multi-turn dialogue training examples, covering system messages, user inputs, assistant responses, and tool calls (if applicable); the difficulty column identifies the difficulty level of each sample, categorized as easy, medium, hard, and adversarial. This version is explicitly labeled as a preview and may have issues, not being the final production release.
数据集概述
- 名称:
mediumish-small-agent-sft-preview-v0.1 - 许可证:
other(未具体说明) - 语言: 英语 (
en) - 任务类别: 文本生成 (
text-generation) - 数据大小: 1K < n < 10K 行(共 6,976 行)
数据集描述
该数据集是通过程序化生成的 ChatML 格式 SFT 数据,专为中小型模型设计,覆盖以下能力:
- 代理工具使用(agentic tool use)
- 反幻觉习惯(anti-hallucination habits)
- 有依据的拒绝(grounded refusal)
- 多步推理(multi-step reasoning)
- 相关认知行为(epistemic behaviors)
领域分布
数据横跨 20 个领域,每个领域字符占比被刻意限制,避免单一领域占据过多 token:
- 可靠性/工具使用
- 圣经研究
- 隐含假设推理
- 高级数学
- 符合文档规范的代码修复
- 规则/政策模拟
- ARC 风格网格谜题
- 状态机
- 逻辑网格
- 引用消解
- 反事实推理
- 及其他更多领域
字符占比示例:
code领域约占 28%- ARC 网格交互领域低于 5%
数据列
| 列名 | 说明 |
|---|---|
chatml |
完整的多轮训练示例(包含 system/user/assistant/tool 轮次) |
difficulty |
难度级别:easy, medium, hard, adversarial |
注意事项
- 当前版本为 预览版,可能存在缺陷,非最终生产版本。




