dev_set_v2_a1_nemotron_bash_withtests_gpt5mini_20260813_031150
收藏资源简介:
该数据集包含 5364 个训练样本,每个样本记录了一次多轮对话(conversations),对话由角色(role)和内容(content)组成。此外,每个样本还包含以下字段:使用的智能体(agent)、模型(model)及模型提供者(model_provider)、日期(date)、所属任务(task)、轮次(episode)、运行ID(run_id)、试验名称(trial_name)、结果(result)、验证器输出(verifier_output)以及追踪来源(trace_source)。数据集可能用于分析不同模型或智能体在特定任务中的对话表现、结果追踪及验证。
This dataset contains 5364 training samples, each recording a multi-turn conversation consisting of role and content. Additionally, each sample includes the following fields: agent used, model and model provider, date, task, episode, run_id, trial_name, result, verifier_output, and trace_source. The dataset may be used for analyzing the conversational performance of different models or agents in specific tasks, as well as result tracking and verification.
数据集概述
基本信息
- 数据集名称:dev_set_v2_a1_nemotron_bash_withtests_gpt5mini_20260813_031150
- 数据集地址:https://huggingface.co/datasets/laion/dev_set_v2_a1_nemotron_bash_withtests_gpt5mini_20260813_031150
- 发布机构:LAION
数据集规模
- 总大小:约 473.9 MB(数据集大小)
- 下载大小:约 403.5 MB
- 数据分割:仅包含训练集(train),共 5,364 条样本
数据字段结构
每条样本包含以下字段:
| 字段名 | 类型 | 说明 |
|---|---|---|
| conversations | 列表 | 对话记录,每条对话包含角色(role)和内容(content)两个子字段 |
| agent | 字符串 | 代理标识 |
| model | 字符串 | 使用的模型名称 |
| model_provider | 字符串 | 模型提供方 |
| date | 字符串 | 日期信息 |
| task | 字符串 | 任务类型 |
| episode | 字符串 | 会话轮次/情节 |
| run_id | 字符串 | 运行标识 |
| trial_name | 字符串 | 试验名称 |
| result | 字符串 | 结果信息 |
| verifier_output | 字符串 | 验证器输出 |
| trace_source | 字符串 | 追踪来源 |
数据特点
- 该数据集为对话式数据,核心字段为
conversations,记录多轮人机对话内容 - 数据集中包含代理(agent)、模型(model) 及模型提供方(model_provider) 等元数据,可用于模型性能溯源
- 包含任务类型(task) 和验证器输出(verifier_output),适用于模型推理能力评估与验证场景
- 数据集名暗示其可能来自Nemotron模型与GPT-5 mini的对比或测试数据,主题涉及Bash命令相关任务(含测试)




