terminal_bench_2_a1_tulu3_sft_personas_math_20260810_081853
收藏资源简介:
该数据集包含多轮对话记录及相关的元数据,用于 AI 代理或模型的评估与分析。每个样本包含一个对话列表(conversations),每条消息具有角色(role)和内容(content);此外还包括代理标识(agent)、模型名称(model)、模型提供者(model_provider)、日期(date)、任务类型(task)、情节编号(episode)、运行ID(run_id)、试验名称(trial_name)、任务结果(result)、验证器输出(verifier_output)以及跟踪来源(trace_source)。数据集中仅包含训练集,共 4495 个样本,总大小约 345 MB,适用于对话系统评估、模型行为分析或验证任务。
This dataset contains multi-turn conversation records and associated metadata for evaluation and analysis of AI agents or models. Each sample includes a list of conversations, where each message has a role and content; additionally, it includes agent identifier, model name, model provider, date, task type, episode number, run ID, trial name, task result, verifier output, and trace source. The dataset contains only a training set with 4495 samples, totaling approximately 345 MB, suitable for dialogue system evaluation, model behavior analysis, or verification tasks.
数据集概述:laion/terminal_bench_2_a1_tulu3_sft_personas_math_20260810_081853
该数据集由 LAION 发布,是 Terminal-Bench 2.0 的一个子集,专注于数学任务的智能体行为数据,基于 Tulu3 SFT 模型生成,并配有 Personas 配置。
数据集结构
特征字段
- conversations (列表): 包含对话轮次,每轮包含:
role(字符串): 消息角色(如用户、助手)content(字符串): 消息内容
- agent (字符串): 智能体标识
- model (字符串): 使用的模型名称
- model_provider (字符串): 模型提供方
- date (字符串): 数据生成日期
- task (字符串): 任务标识
- episode (字符串): 回合编号
- run_id (字符串): 运行标识
- trial_name (字符串): 试验名称
- result (字符串): 任务结果
- verifier_output (字符串): 验证器输出
- trace_source (字符串): 追踪来源
数据划分
- 划分名称: train
- 样本数量: 4,495 条
- 数据大小: 345,958,856 字节(约 330 MB)
- 下载大小: 284,757,293 字节(约 272 MB)
配置信息
- 配置名称: default
- 数据文件路径:
data/train-*(分片存储)
用途说明
该数据集适用于智能体行为分析、数学任务求解过程研究、对话系统评估等场景,尤其关注终端环境中基于 Tulu3 SFT 模型的自主智能体表现。




