swebench_verified_random_100_folders_a3_rl_DCAgent_exp_rpt_pymethods2test_v3_10_8Bc8b3394d
收藏资源简介:
该数据集记录了AI代理与语言模型之间的多轮对话及其相关元数据,包含以下字段:对话历史(conversations,每个对话包含角色role和内容content)、代理标识(agent)、模型名称(model)、模型提供者(model_provider)、日期(date)、任务描述(task)、回合编号(episode)、运行ID(run_id)、试验名称(trial_name)、结果(result)、验证器输出(verifier_output)以及追踪来源(trace_source)。数据集共包含5055个训练样本,总大小约1GB,适用于研究代理与模型的交互行为、任务执行效果、结果验证等场景。
This dataset records multi-turn conversations between AI agents and language models along with their associated metadata. It includes the following fields: conversation history (conversations, each containing role and content), agent identifier (agent), model name (model), model provider (model_provider), date (date), task description (task), episode number (episode), run ID (run_id), trial name (trial_name), result (result), verifier output (verifier_output), and trace source (trace_source). The dataset contains a total of 5055 training samples, with a total size of approximately 1GB, suitable for studying agent-model interaction behavior, task execution effectiveness, and result verification.
数据集概述:laion/swebench_verified_random_100_folders_a3_rl_DCAgent_exp_rpt_pymethods2test_v3_10_8Bc8b3394d
基本信息
- 数据集名称:laion/swebench_verified_random_100_folders_a3_rl_DCAgent_exp_rpt_pymethods2test_v3_10_8Bc8b3394d
- 数据集地址:https://huggingface.co/datasets/laion/swebench_verified_random_100_folders_a3_rl_DCAgent_exp_rpt_pymethods2test_v3_10_8Bc8b3394d
- 数据集大小:约 10亿字节(1,001,766,687 字节)
- 下载大小:约 4.7亿字节(469,748,185 字节)
数据划分
该数据集仅包含一个划分:
- train(训练集):共 5,055 个样本
数据特征(列字段)
数据集包含以下字段:
| 字段名 | 类型 | 说明 |
|---|---|---|
| conversations | 列表(包含 role 和 content 两个子字段,均为字符串类型) | 对话记录,包含角色和内容 |
| agent | 字符串 | 智能体名称 |
| model | 字符串 | 模型名称 |
| model_provider | 字符串 | 模型提供商 |
| date | 字符串 | 日期 |
| task | 字符串 | 任务描述 |
| episode | 字符串 | 回合/片段编号 |
| run_id | 字符串 | 运行标识符 |
| trial_name | 字符串 | 试验名称 |
| result | 字符串 | 结果 |
| verifier_output | 字符串 | 验证器输出 |
| trace_source | 字符串 | 轨迹来源 |
数据格式
- 配置文件名称:default
- 数据文件路径:data/train-*(多个分片文件)
数据集内容推测
从数据集名称和字段结构来看,该数据集似乎与SWE-bench验证任务相关,涉及智能体(agent)在强化学习(RL)环境下的交互轨迹数据,可能记录了智能体在代码生成、测试或软件工程任务中的对话历史、模型响应、验证结果等信息。




