dev_set_v2_a1_e2egit_20260820_135118
收藏资源简介:
该数据集包含多轮对话记录及其相关元数据。每条数据包括对话历史(conversations,由角色和内容组成)、代理标识(agent)、使用的模型名称(model)及提供商(model_provider)、记录日期(date)、任务描述(task)、实验回合(episode)、运行ID(run_id)、试验名称(trial_name)、最终结果(result)、验证器输出(verifier_output)以及跟踪来源(trace_source)。训练集共有5134个样本,适用于对话系统、AI代理行为分析、模型评估和验证等场景。
This dataset contains multi-turn conversation records and their associated metadata. Each entry includes conversation history (conversations, consisting of roles and content), agent identifier (agent), model name (model) and provider (model_provider), recording date (date), task description (task), experiment episode (episode), run ID (run_id), trial name (trial_name), final result (result), verifier output (verifier_output), and trace source (trace_source). The training set contains 5134 samples, suitable for dialogue systems, AI agent behavior analysis, model evaluation and validation, etc.
数据集概述
- 名称:
laion/dev_set_v2_a1_e2egit_20260820_135118 - 链接: 数据集详情页
基本信息
- 数据集大小: 612,194,412 字节(约 584 MB)
- 下载大小: 529,271,574 字节(约 505 MB)
- 数据划分: 仅包含
train划分- 样本数量: 5,134
- 数据字节数: 612,194,412
数据特征
数据集包含以下字段:
- conversations: 对话列表,每个对话包含:
role: 角色(字符串)content: 内容(字符串)
- agent: 代理标识(字符串)
- model: 模型名称(字符串)
- model_provider: 模型提供方(字符串)
- date: 日期(字符串)
- task: 任务类型(字符串)
- episode: 会话轮次(字符串)
- run_id: 运行标识(字符串)
- trial_name: 试验名称(字符串)
- result: 结果(字符串)
- verifier_output: 验证器输出(字符串)
- trace_source: 追踪来源(字符串)
数据配置
- 配置名称:
default - 数据文件:
data/train-*(通配符匹配多个文件)




