sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13
收藏资源简介:
该数据集包含对 `laion/sft-repro-thinking-step630-nemotron-terminal-step1888` 模型在 OpenThoughts-TBLite 基准上进行 300 次试验的完整评估结果。所有试验均使用 Terminus-2 框架,尝试次数为 3,温度 0.6,上下文窗口为 24,576 输入令牌 + 8,192 输出令牌。数据集中包含每个试验的轨迹、验证器输出、FineStore 表、聚合统计文件以及启动记录。敏感信息(如终端令牌和凭证)已被替换为明确的删除标记,替换记录保存在 `redaction-report.json` 中。可用的评估指标包括尝试/完成次数(300/300)、可评分率(86.33%)、总奖励(0.1449)、可评分试验平均奖励(0.1678)等。该数据集可用于复现分析、评估方法研究或作为 agentic 评估的基准。
This dataset contains the complete evaluation results of 300 trials on the OpenThoughts-TBLite benchmark for the `laion/sft-repro-thinking-step630-nemotron-terminal-step1888` model. All trials were conducted using the Terminus-2 framework, with 3 attempts, temperature 0.6, and a context window of 24,576 input tokens + 8,192 output tokens. The dataset includes trajectories, validator outputs, FineStore tables, aggregated statistics files, and launch records for each trial. Sensitive information (such as terminal tokens and credentials) has been replaced with explicit deletion markers, with replacement records saved in `redaction-report.json`. Available evaluation metrics include attempts/completions (300/300), scoreable rate (86.33%), total reward (0.1449), average reward per scoreable trial (0.1678), and others. This dataset can be used for reproducibility analysis, evaluation method research, or as a benchmark for agentic evaluations.
数据集概述:laion/sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-2026-08-13
数据集简介
该数据集是 Nemotron Terminal SFT 复现评估工件(Harbor 工件树),用于记录 300 次 OpenThoughts-TBLite 评估试验的完整结果。被评估的检查点模型为 laion/sft-repro-thinking-step630-nemotron-terminal-step1888,该模型基于 Grug 阶段二思维检查点,在 Nemotron Terminal 语料上训练了 1,888 步。
评估结果
| 指标 | 数值 |
|---|---|
| 尝试 / 完成次数 | 300 / 300 |
| 可进行 Verifier 评分 | 259 (86.33%) |
| 所有尝试的聚合奖励 | 0.1449016619 |
| 可评分试验的平均奖励 | 0.1678397628 |
| 满分奖励试验数 | 39 |
| 非零奖励试验数 | 49 |
| 上下文管理基础设施错误 | 39 |
内容与结构
- 存档内容:所有试验结果、轨迹、验证器输出、FineStore 表、聚合文件以及启动记录。
- 安全处理:签名 Iris 端点令牌和凭证形状的任务夹具内容在发布前被替换为显式的删除标记。
- 审计文件:
redaction-report.json记录了每次替换的形状和文件位置,但不保留原始机密文本。
溯源信息
- 运行标识:
20260812-200420-sft-repro-thinking-step630-nemotron-terminal-step1888-openthoughts-tblite-terminus2-control-a40b - 基础数据集:
open-thoughts/OpenThoughts-TBLite@main - 评估框架:Terminus-2
- 每任务尝试次数:3
- 采样温度:0.6
- 上下文长度:24,576 输入 + 8,192 输出 tokens
- 组件版本:Marin
505d4d799bdf9777aedc03945d0ae0f050460013、Harbor24e5e67ac93d21c1a2202d6906f6093e0a57b86c、Evalchemy5ef8b1604a535d04798f3424daa3fe4facd9bf32
许可证与任务
- 许可证:Apache-2.0
- 任务类别:文本生成(text-generation)
- 标签:agentic-evaluation、harbor、terminus-2、openthoughts-tblite
附加资源
精确的训练启动器、导出驱动、模型/评估配置、启动记录及分析源码可参阅 复现源码 gist。




