terminal_bench_2_a1_swegym_openhands_20260808_162055
收藏资源简介:
该数据集包含多轮对话记录,每条记录由 conversations 字段(对话轮次列表,每轮标注角色和内容)、智能体名称(agent)、模型名称(model)及提供商(model_provider)、日期(date)、任务(task)、回合编号(episode)、运行ID(run_id)、试验名称(trial_name)、结果(result)、验证器输出(verifier_output)和追踪来源(trace_source)组成。训练集共有 6719 个样本,总数据量约 579 MB。数据集可用于对话系统开发、智能体行为分析、模型评估与追踪等任务。
This dataset contains multi-turn conversation records. Each record consists of the conversations field (a list of dialogue turns, each turn annotated with role and content), agent name (agent), model name (model) and provider (model_provider), date (date), task (task), episode number (episode), run ID (run_id), trial name (trial_name), result (result), verifier output (verifier_output), and trace source (trace_source). The training set has 6719 samples, with a total data size of approximately 579 MB. The dataset can be used for dialogue system development, agent behavior analysis, model evaluation and tracking, etc.
数据集概述
该数据集名为 terminal_bench_2_a1_swegym_openhands_20260808_162055,由 LAION 发布,旨在记录和评估 AI 智能体在终端环境中的任务执行轨迹。
数据集内容
数据集包含 6,719 条训练样本,每条样本记录了 AI 智能体在一次任务执行过程中的完整对话历史、任务信息与最终结果。具体字段如下:
| 字段名 | 说明 |
|---|---|
conversations |
对话历史,包含角色(role)和内容(content)两部分 |
agent |
使用的智能体标识 |
model |
底层模型名称 |
model_provider |
模型提供方 |
date |
执行日期 |
task |
任务描述 |
episode |
回合编号 |
run_id |
运行标识 |
trial_name |
试验名称 |
result |
任务结果 |
verifier_output |
验证器输出 |
trace_source |
轨迹来源 |
数据规模
- 数据集总大小:约 579.8 MB(未压缩)
- 下载大小:约 461.1 MB(压缩后)
- 样本数量:6,719 条(训练集)
数据划分
该数据集仅包含一个 train 划分,数据文件存储在 data/train-* 路径下。
适用场景
该数据集适用于研究 AI 智能体在终端命令执行、软件开发任务(如 SWE-bench 类问题)中的行为模式、推理过程及任务成功率,可用于训练、评估或分析终端智能体的性能。




