terminal_bench_2_a1_stack_dockerfile_20260820_214936
收藏资源简介:
该数据集包含6117个训练样本,每个样本记录一次智能体对话交互过程。数据字段包括:对话历史(conversations,由角色和内容组成)、智能体名称(agent)、模型名称(model)、模型提供者(model_provider)、日期(date)、任务类型(task)、对话轮次(episode)、运行ID(run_id)、试验名称(trial_name)、最终结果(result)、验证器输出(verifier_output)以及追踪来源(trace_source)。数据集适用于研究多轮对话、智能体行为分析、模型评估或对话系统训练等任务。
This dataset contains 6117 training samples, each recording an agent dialogue interaction process. The data fields include: conversation history (conversations, consisting of roles and content), agent name (agent), model name (model), model provider (model_provider), date (date), task type (task), episode (episode), run ID (run_id), trial name (trial_name), final result (result), verifier output (verifier_output), and trace source (trace_source). The dataset is suitable for studying multi-turn dialogue, agent behavior analysis, model evaluation, or dialogue system training.
数据集概述
该数据集名为 terminal_bench_2_a1_stack_dockerfile_20260820_214936,由 LAION 组织发布,托管于 Hugging Face 平台。数据集主要面向终端(Terminal)相关的基准测试场景,集中于 Dockerfile 构建任务,数据生成日期为 2026 年 8 月 20 日。
数据集结构
- 特征字段(Features):
conversations(对话记录):包含role(角色)和content(内容)两个子字段,属于列表类型。agent(智能体名称):字符串类型。model(模型名称):字符串类型。model_provider(模型提供商):字符串类型。date(日期):字符串类型。task(任务描述):字符串类型。episode(会话轮次):字符串类型。run_id(运行编号):字符串类型。trial_name(试验名称):字符串类型。result(执行结果):字符串类型。verifier_output(验证器输出):字符串类型。trace_source(追踪来源):字符串类型。
数据划分与规模
- 数据划分(Splits):
train(训练集):包含 6117 个样本,占用存储空间约 500 MB(500,106,175 字节)。
- 数据集总大小:约 500 MB(500,106,175 字节)。
- 下载大小:约 407 MB(407,464,817 字节)。
配置文件
- 配置名称:
default(默认配置)。 - 数据文件路径:
data/train-*(对应训练集分片文件)。
数据用途
该数据集面向终端环境下的基准测试,重点覆盖 Dockerfile 相关任务,可能用于评估或训练智能体(agent)在终端操作、代码生成、构建配置等场景中的表现。每条样本包含完整的对话记录、执行结果及验证信息,适合用于构建或评测终端代理模型。




