eval-penfever_stageD-thinkbudget-80-8B_DCAgent_dev_set_v2-traces
收藏资源简介:
该数据集包含用于对话和任务执行场景的结构化数据,主要特征包括多轮对话记录(每条消息包含内容和角色信息)、代理标识、模型信息(如模型名称和提供方)、日期、任务类型、情节编号、运行ID、试验名称、执行结果、验证器输出以及数据追踪来源。数据集规模为1274个训练样本,适用于多轮对话分析、代理行为评估、任务完成度验证或模型输出比较等应用场景。数据以结构化字段形式组织,支持对对话流程、模型性能和任务结果的深入分析。
This dataset contains structured data for dialogue and task execution scenarios, with key features including multi-turn dialogue records (each message contains content and role information), agent identifiers, model information (such as model name and provider), date, task type, episode number, run ID, trial name, execution results, validator output, and data tracking sources. The dataset comprises 1274 training samples and is suitable for applications such as multi-turn dialogue analysis, agent behavior evaluation, task completion verification, or model output comparison. The data is organized in structured fields, enabling in-depth analysis of dialogue flow, model performance, and task outcomes.
- 数据集名称:eval-penfever_stageD-thinkbudget-80-8B_DCAgent_dev_set_v2-traces
- 来源:LAION 团队
- 数据集大小:约 95.45 MB(下载大小约 86.82 MB)
- 数据量:1274 条样本
- 数据划分:仅包含训练集(train)
- 特征字段:
conversations:对话记录列表,每条包含content(字符串)和role(字符串)agent:代理标识(字符串)model:使用的模型名称(字符串)model_provider:模型提供商(字符串)date:日期(字符串)task:任务描述(字符串)episode:轮次标识(字符串)run_id:运行ID(字符串)trial_name:试验名称(字符串)result:结果(字符串)verifier_output:验证器输出(字符串)trace_source:追踪来源(字符串)
- 数据用途:该数据集主要用于评估或追踪特定代理(DCAgent)在给定任务上的表现,包含对话历史、模型信息、运行元数据及结果验证信息




