dev_set_v2_a1_magicoder_20260805_125433
收藏资源简介:
该数据集包含多轮对话记录,每条样本由对话内容(conversations,包括角色和文本)、智能体(agent)、模型名称(model)、模型提供商(model_provider)、日期(date)、任务(task)、轮次(episode)、运行ID(run_id)、试验名称(trial_name)、结果(result)、验证器输出(verifier_output)以及追踪来源(trace_source)组成。数据集共有3293个训练样本,适用于对话系统评估、模型性能比较、任务导向对话分析等场景。
This dataset contains multi-turn conversation records. Each sample consists of conversation content (conversations, including role and text), agent, model name, model provider, date, task, episode, run ID, trial name, result, verifier output, and trace source. The dataset has a total of 3293 training samples and is suitable for dialogue system evaluation, model performance comparison, task-oriented dialogue analysis, etc.
数据集概述
该数据集名为 dev_set_v2_a1_magicoder_20260805_125433,由 LAION 组织提供,托管于 Hugging Face 平台。数据集包含 3,293 条训练样本,总大小为 339,502,907 字节(约 324 MB),下载大小为 303,070,485 字节(约 289 MB)。
数据特征
数据集包含以下字段:
| 字段名 | 类型 | 说明 |
|---|---|---|
conversations |
列表 | 对话内容,包含 role(角色,字符串)和 content(内容,字符串)两个子字段 |
agent |
字符串 | 代理标识 |
model |
字符串 | 使用的模型名称 |
model_provider |
字符串 | 模型提供商 |
date |
字符串 | 日期信息 |
task |
字符串 | 任务类型 |
episode |
字符串 | 轮次信息 |
run_id |
字符串 | 运行标识 |
trial_name |
字符串 | 试验名称 |
result |
字符串 | 结果信息 |
verifier_output |
字符串 | 验证器输出 |
trace_source |
字符串 | 轨迹来源 |
数据集划分
该数据集仅包含一个划分:
- train:包含 3,293 个样本,大小为 339,502,907 字节。
数据文件存储在 data/train-* 路径下,默认配置名为 default。
数据规模概览
- 样本数量:3,293 条
- 数据集总大小:约 324 MB
- 下载大小:约 289 MB
该数据集可能用于多轮对话、智能体行为追踪或模型评估等相关任务的研究与开发。




