eval-laion_symclip-30-8B_DCAgent2_terminal_bench_2-traces
收藏资源简介:
该数据集是一个用于记录和分析智能体或模型在任务导向交互中表现的数据集,包含多轮对话内容,每个样本涵盖对话历史、参与角色、使用的模型及提供者、执行日期、任务类型、运行标识、试验名称、任务结果以及验证输出等字段。数据以结构化形式存储,适用于对话系统评估、智能体行为分析、模型性能比较等应用场景。数据规模为1168个训练样本,总大小约92MB。
This dataset is designed for recording and analyzing the performance of AI Agents or models during task-oriented interactions. It contains multi-turn conversation content, with each sample covering fields such as conversation history, participating roles, the models used and their providers, execution date, task type, run ID, experiment name, task results, and validation outputs. The data is stored in a structured format, and is applicable to scenarios including dialogue system evaluation, AI Agent behavior analysis, and model performance comparison. The dataset has 1168 training samples with a total size of approximately 92 MB.
此数据集存储于 Hugging Face,旨在提供特定场景下 AI Agent 的交互轨迹数据。
数据集基本信息
- 数据集名称:eval-laion_symclip-30-8B_DCAgent2_terminal_bench_2-traces
- 数据集大小:约 87.9 MB(下载大小约 71.4 MB)
- 样本数量:1,168 条
- 数据拆分:仅包含一个训练集(train)
数据特征(Fields)
每条数据包含以下字段:
conversations:对话历史,是一个列表,每个列表元素包含content(字符串,对话内容)和role(字符串,发言角色)。agent:使用的 Agent 名称(字符串)。model:使用的模型名称(字符串)。model_provider:模型提供方(字符串)。date:数据记录日期(字符串)。task:执行的任务描述(字符串)。episode:任务回合编号(字符串)。run_id:运行 ID(字符串)。trial_name:试验名称(字符串)。result:任务结果(字符串)。verifier_output:验证器输出(字符串)。trace_source:轨迹来源(字符串)。
数据用途
该数据集可用于分析或微调 AI Agent 在终端任务上的表现,包括对话记录、任务结果及验证信息。




