dev_set_v2_a3_rl_laion_nemotron_gym_knowledge_web_search_mcqa_25_8B_20260825_134037
收藏资源简介:
该数据集包含多轮对话记录,每个样本由对话内容(conversations,包含角色和文本)、使用的智能体(agent)、模型(model)及其提供商(model_provider)、日期(date)、任务(task)、回合(episode)、运行ID(run_id)、试验名称(trial_name)、结果(result)、验证器输出(verifier_output)和追踪来源(trace_source)等字段组成。数据集共有3427个训练样本,总大小约252MB。适用于对话系统训练、智能体行为分析、模型评估与验证等任务。
This dataset contains multi-turn dialogue records. Each sample consists of conversations (including roles and text), the agent used, the model and its provider, date, task, episode, run ID, trial name, result, verifier output, and trace source. The dataset has 3427 training samples with a total size of approximately 252 MB. It is suitable for tasks such as dialogue system training, agent behavior analysis, model evaluation, and validation.
数据集概述
该数据集名为 laion/dev_set_v2_a3_rl_laion_nemotron_gym_knowledge_web_search_mcqa_25_8B_20260825_134037,由 LAION 组织发布,主要用于强化学习、知识问答、网络搜索和多选题(MCQA)等场景的训练与评估。
基本信息
- 数据集规模:训练集包含 3,427 个样本,总大小约 252.4 MB(下载大小约 217.7 MB)。
- 数据划分:仅提供
train划分,数据存储在data/train-*路径下。
数据特征
每个样本包含以下字段:
| 字段名 | 类型 | 说明 |
|---|---|---|
conversations |
列表 | 对话记录,每项包含 role(角色)和 content(内容),均为字符串类型 |
agent |
字符串 | 智能体标识 |
model |
字符串 | 使用的模型名称(如 8B 模型) |
model_provider |
字符串 | 模型提供方 |
date |
字符串 | 日期信息 |
task |
字符串 | 任务类型(如知识问答、网络搜索、MCQA 等) |
episode |
字符串 | 回合/轮次标识 |
run_id |
字符串 | 运行 ID |
trial_name |
字符串 | 试验名称 |
result |
字符串 | 任务结果 |
verifier_output |
字符串 | 验证器输出 |
trace_source |
字符串 | 轨迹来源 |
用途说明
该数据集可能适用于:
- 强化学习(RL)训练与评估
- 知识问答(knowledge)任务
- 网络搜索(web search)场景
- 多选题(MCQA)测试
- 对话系统开发与验证




