terminal_bench_2_tasktrove_dq_tulu3_personas_math_step13_30b_a3b_20260730_054033
收藏资源简介:
该数据集名为 terminal_bench_2_tasktrove_dq_tulu3_personas_math_step13_30b_a3b,是一个从Iris强化学习运行中导出的OpenCode智能体轨迹数据集。数据集包含了智能体在数学任务上的交互对话记录,覆盖了完整的试验过程。每条数据包含智能体与环境的对话(conversations字段,包括角色和内容)、智能体标识(agent)、使用的模型(model及model_provider)、日期(date)、任务描述(task)、试验编号(episode)、运行ID(run_id)、试验名称(trial_name)、结果(result)、指令(instruction)以及验证器输出(verifier_output)。数据集共包含15,651个训练样本,总大小约118MB。所有生成了result.json的试验均被收录,缺失的89个试验因未完成可评分的片段而未包含。该数据集适用于强化学习、智能体行为分析、数学推理等任务的研究。
This dataset is named terminal_bench_2_tasktrove_dq_tulu3_personas_math_step13_30b_a3b, an OpenCode agent trajectory dataset derived from an Iris reinforcement learning run. It contains interactive dialogue records of agents on mathematical tasks, covering the complete trial process. Each entry includes conversations between the agent and environment (conversations field with role and content), agent identifier (agent), model used (model and model_provider), date, task description (task), trial number (episode), run ID (run_id), trial name (trial_name), result, instruction, and verifier output (verifier_output). The dataset contains 15,651 training samples with a total size of approximately 118MB. All trials that generated result.json are included; the missing 89 trials were not included because they did not complete a scorable segment. This dataset is suitable for research in reinforcement learning, agent behavior analysis, mathematical reasoning, and other tasks.
数据集概述
该数据集名为 terminal_bench_2_tasktrove_dq_tulu3_personas_math_step13_30b_a3b,来源于一次名为 rl-tasktrove-dq-sweep-30b-qwen3-coder-30-20260727-143750-b2bcd7 的 Iris 强化学习运行,从该运行的 Harbor 的 trace_jobs 工件中导出(每个 trial 的最后一个 episode)。
数据规模
| 项目 | 数量 |
|---|---|
| trial 目录总数 | 15,740 |
具有 result.json 的 trial 数 |
15,651 |
| 数据集行数 | 15,651 |
覆盖完整:所有 trial 目录均被枚举,每个产生 result.json 的 trial 均包含在内。89 个没有 result.json 的 trial 因未完成可计分 episode 而未贡献行。
注意:该数据集早期版本仅含 255 行,基于部分本地镜像构建,现已被完整导出版本取代。
数据特征
每条记录包含以下字段:
- conversations:对话列表,包含
content(字符串)和role(字符串) - agent:智能体标识(字符串)
- model:模型名称(字符串)
- model_provider:模型提供方(字符串)
- date:日期(字符串)
- task:任务(字符串)
- episode:episode 编号(字符串)
- run_id:运行 ID(字符串)
- trial_name:trial 名称(字符串)
- result:结果(字符串)
- instruction:指令(字符串)
- verifier_output:验证器输出(字符串)
数据划分
- 训练集 (train):15,651 条样本,占用 118,065,300 字节(约 112.6 MB)
- 数据集总大小:118,065,300 字节
- 下载大小:113,630,521 字节
数据文件路径为 data/train-*。




