eval-laion_seqnorm-tis-shaped-40-8B_DCAgent2_terminal_bench_2-traces
收藏资源简介:
该数据集包含1053个对话交互样本,主要用于任务导向的对话系统评估或智能代理测试。每个样本记录完整的对话过程(conversations字段,包含content和role),并附带丰富的元数据:包括使用的代理类型(agent)、模型名称(model)及提供商(model_provider)、任务类型(task)、交互轮次(episode)、运行标识(run_id)、试验名称(trial_name)、任务结果(result)、验证器输出(verifier_output)以及数据来源追踪(trace_source)。数据还包含时间戳(date)信息。数据集结构表明其适用于分析多轮对话性能、评估不同模型/代理在特定任务上的表现,以及研究对话系统的可追溯性。
This dataset contains 1053 dialogue interaction samples, primarily used for task-oriented dialogue system evaluation or intelligent agent testing. Each sample records the complete dialogue process (conversations field, including content and role) and is accompanied by rich metadata: including the agent type used (agent), model name (model) and provider (model_provider), task type (task), interaction rounds (episode), run identifier (run_id), trial name (trial_name), task result (result), verifier output (verifier_output), and data source tracking (trace_source). The data also includes timestamp (date) information. The dataset structure indicates its suitability for analyzing multi-turn dialogue performance, evaluating the performance of different models/agents on specific tasks, and studying the traceability of dialogue systems.
- 名称: eval-laion_seqnorm-tis-shaped-40-8B_DCAgent2_terminal_bench_2-traces
- 地址: https://huggingface.co/datasets/laion/eval-laion_seqnorm-tis-shaped-40-8B_DCAgent2_terminal_bench_2-traces
- 数据类型: 包含对话、代理、模型、日期、任务、回合、运行ID、试验名称、结果、验证器输出和轨迹来源等字段。
- 特征:
- conversations: 由多个消息构成的列表,每个消息包含 content(文本内容)和 role(角色)字段。
- agent: 字符串,记录代理名称。
- model: 字符串,记录模型名称。
- model_provider: 字符串,记录模型提供者。
- date: 字符串,记录日期。
- task: 字符串,记录任务名称。
- episode: 字符串,记录回合信息。
- run_id: 字符串,记录运行标识。
- trial_name: 字符串,记录试验名称。
- result: 字符串,记录结果。
- verifier_output: 字符串,记录验证器输出。
- trace_source: 字符串,记录轨迹来源。
- 数据划分:
- 训练集(train):包含 1053 条样本,占用 81,801,268 字节。
- 下载大小: 65,669,642 字节
- 数据集总大小: 81,801,268 字节
- 配置文件: 默认配置(default),数据文件路径为 data/train-*。




