DCAgent2/terminal_bench_2_Qwen3_Coder_480B_A35B_Instruct_FP8_20260429_192429
收藏资源简介:
该数据集是一个包含多轮对话和任务执行记录的数据集,主要用于模型评估和代理交互研究。数据集包含267个训练示例,每个示例具有多个特征:conversations(对话列表,包括内容和角色)、agent(代理标识)、model(模型名称)、model_provider(模型提供者)、date(日期)、task(任务类型)、episode(剧集编号)、run_id(运行ID)、trial_name(试验名称)、result(执行结果)和verifier_output(验证器输出)。数据以训练分割形式提供,总大小约23MB,适用于自然语言处理、对话系统和人工智能代理的评估任务。
This dataset is a collection of multi-turn conversations and task execution records, primarily designed for model evaluation and agent interaction research. It contains 267 training examples, each with multiple features: conversations (a list of dialogues including content and role), agent (agent identifier), model (model name), model_provider (model provider), date (date of execution), task (task type), episode (episode number), run_id (run ID), trial_name (trial name), result (execution result), and verifier_output (verifier output). The data is provided in a train split format, with a total size of approximately 23MB, and is suitable for natural language processing, dialogue systems, and evaluation tasks for AI agents.




