eval-fsr-a1-tulu3-sft-personas-math-swe-r474-traces
收藏资源简介:
该数据集记录了多轮对话交互实验的详细信息,包含1130个训练样本。每个样本的核心内容为多轮对话(conversations),其中每轮包含消息内容(content)和发言者角色(role)。此外,样本还包含丰富的元数据:执行对话的代理(agent)、使用的模型(model)及其提供商(model_provider)、实验日期(date)、任务类型(task)、实验轮次(episode)、运行标识(run_id)、试验名称(trial_name)、实验结果(result)、验证器输出(verifier_output)以及数据来源(trace_source)。数据集适用于分析AI代理在对话任务中的表现、评估模型交互能力或研究多轮对话的实验设计。
This dataset records detailed information from multi-round dialogue interaction experiments, containing 1130 training samples. Each samples core content consists of multi-round conversations, where each round includes message content and speaker role. Additionally, samples include rich metadata: the agent conducting the dialogue, the model used and its provider, experiment date, task type, episode number, run identifier, trial name, experiment result, verifier output, and data source. The dataset is suitable for analyzing AI agent performance in dialogue tasks, evaluating model interaction capabilities, or studying experimental designs for multi-round conversations.




