eval-fsr-a1-go-browse-wa-swe-r444-traces
收藏资源简介:
该数据集记录了AI代理或模型在特定任务上的交互轨迹与执行结果。数据集的核心内容是多轮对话记录(包含每轮的内容和发言角色),并关联了丰富的元数据,包括使用的代理类型、模型名称、模型提供方、执行日期、任务类型、对话片段标识、运行ID、试验名称、任务执行结果、验证器输出以及数据追踪来源。数据集包含3226个样本,仅提供训练集划分,总大小约为515MB。从数据结构推断,该数据集适用于分析AI系统的对话行为、评估代理在不同任务上的性能、研究强化学习中的交互轨迹,或用于训练和验证对话系统及任务导向型AI模型。
This dataset records the interaction trajectories and execution results of AI agents or models on specific tasks. The core content consists of multi-turn dialogue records (including the content of each turn and the speakers role), associated with rich metadata such as the agent type used, model name, model provider, execution date, task type, dialogue segment identifier, run ID, experiment name, task execution result, validator output, and data source tracking. The dataset contains 3226 samples, with only a training set split provided, and has a total size of approximately 515MB. Based on the data structure, it is suitable for analyzing the dialogue behavior of AI systems, evaluating agent performance across different tasks, studying interaction trajectories in reinforcement learning, or for training and validating dialogue systems and task-oriented AI models.




