eval-laion_loopshape-30-8B_DCAgent2_swebench-verified-random-100-folders-traces
收藏资源简介:
该数据集记录了AI代理或模型在特定任务上的执行过程与结果,包含1847个训练样本。每个样本由多个字段构成:conversations字段以列表形式存储多轮对话内容,包括content(文本内容)和role(角色);其他字段如agent(代理标识)、model(模型名称)、model_provider(模型提供商)、date(日期)、task(任务类型)、episode(情节标识)、run_id(运行ID)、trial_name(试验名称)、result(执行结果)、verifier_output(验证输出)和trace_source(追踪来源),共同描述了实验的配置、执行轨迹及验证信息。数据集适用于AI代理评估、任务导向对话分析、模型性能比较等场景,支持对多模态实验运行数据的结构化研究。
This dataset records the execution processes and results of AI agents or models on specific tasks, containing 1847 training samples. Each sample consists of multiple fields: the conversations field stores multi-turn dialogue content in a list format, including content (text content) and role (role); other fields such as agent (agent identifier), model (model name), model_provider (model provider), date (date), task (task type), episode (episode identifier), run_id (run ID), trial_name (trial name), result (execution result), verifier_output (verifier output), and trace_source (trace source) collectively describe the experimental configuration, execution trajectory, and verification information. The dataset is suitable for scenarios such as AI agent evaluation, task-oriented dialogue analysis, and model performance comparison, supporting structured research on multimodal experimental run data.
- 数据集名称: eval-laion_loopshape-30-8B_DCAgent2_swebench-verified-random-100-folders-traces
- 来源: Hugging Face Datasets (laion/eval-laion_loopshape-30-8B_DCAgent2_swebench-verified-random-100-folders-traces)
- 任务类型: 未明确指定,但包含对话、代理、模型、任务等字段,推测为多轮对话或代理行为评估任务
- 数据规模: 训练集包含 1,847 个样本,总数据大小约 306.9 MB,下载大小约 172.0 MB
- 特征字段:
conversations: 对话列表,每条对话包含content(字符串类型)和role(字符串类型)agent: 代理名称(字符串)model: 模型名称(字符串)model_provider: 模型提供者(字符串)date: 日期(字符串)task: 任务描述(字符串)episode: 轮次(字符串)run_id: 运行 ID(字符串)trial_name: 试验名称(字符串)result: 结果(字符串)verifier_output: 验证器输出(字符串)trace_source: 追踪来源(字符串)
- 数据划分: 只有训练集(train),数据文件路径为
data/train-*




