swebench_verified_random_100_folders_g1_a1_top16_32b_step3600_20260819_162251
收藏资源简介:
该数据集包含5122个训练样本,总大小约994MB。每条样本记录了一个多轮对话的完整交互过程,包含对话内容(角色与文本)、智能体名称、使用的模型名称及提供商、对话日期、任务名称、对话轮次(episode)、运行ID(run_id)、试验名称(trial_name)、最终结果(result)、验证器输出(verifier_output)以及追踪来源(trace_source)。该数据集适用于对话智能体训练、模型评估、任务执行追踪与结果验证等场景。
This dataset contains 5122 training samples with a total size of approximately 994MB. Each sample records the complete interaction process of a multi-turn dialogue, including dialogue content (roles and text), agent name, model name and provider used, dialogue date, task name, dialogue episode, run ID, trial name, final result, verifier output, and trace source. This dataset is suitable for scenarios such as dialogue agent training, model evaluation, task execution tracking, and result verification.
数据集详情总结
基本信息
- 数据集名称:
laion/swebench_verified_random_100_folders_g1_a1_top16_32b_step3600_20260819_162251 - 数据集地址:https://huggingface.co/datasets/laion/swebench_verified_random_100_folders_g1_a1_top16_32b_step3600_20260819_162251
数据集规模
- 总大小:约 994 MB(字节数 994,332,409)
- 下载大小:约 436 MB(字节数 436,501,549)
- 数据划分:仅包含
train划分 - 样本数量:5,122 条样本(train 划分)
数据特征(Features)
该数据集包含以下字段:
| 字段名 | 数据类型 | 说明 |
|---|---|---|
conversations |
列表(每项含 role 和 content 字段,均为字符串) |
对话内容,记录交互过程中的角色与消息内容 |
agent |
字符串 | 智能体标识 |
model |
字符串 | 使用的模型名称 |
model_provider |
字符串 | 模型提供方 |
date |
字符串 | 日期信息 |
task |
字符串 | 任务描述 |
episode |
字符串 | 回合/轮次编号 |
run_id |
字符串 | 运行标识 |
trial_name |
字符串 | 试验名称 |
result |
字符串 | 结果信息 |
verifier_output |
字符串 | 验证器输出 |
trace_source |
字符串 | 轨迹来源 |
配置信息
- 配置名称:
default - 数据文件路径:
data/train-*(支持通配符匹配多个 shard 文件)
内容概述
该数据集属于 SWE-bench 验证场景的一个子集,其命名包含随机选择 100 个文件夹、使用特定模型(32B 参数规模,step 3600)生成等元信息。数据集以对话形式存储智能体执行任务过程中的交互记录,并附带了结果、验证器输出、轨迹来源等多维元数据,适用于研究智能体在软件开发任务(SWE-bench)中的行为表现与结果分析。




