huangjh16/BenchTrace
收藏资源简介:
BenchTrace是一个用于评估大型语言模型(LLM)代理自我进化能力的基准测试集。它提供了一个包含1,821个带注释事件的快照-反思数据集,涵盖六个代理任务:ALFWorld具身家庭任务、BabyAI网格世界导航、捆绑式网络购物、ScienceWorld交互式科学、Jericho文本冒险游戏和团体旅行规划。数据集包括两个评估套件:反思评估和进化评估。每个事件包含唯一标识符、环境/游戏实例、任务类型、代理模型、成功状态、快照(用于反思的轨迹上下文)以及带注释的失败实例(包括检测、定位和诊断标签)。数据来源于多个代理环境,遵循各自的许可证,原始注释基于CC-BY-4.0发布。
BenchTrace is a benchmark for evaluating the self-evolution ability of LLM agents. It provides a snapshot–reflection dataset of 1,821 annotated episodes spanning six agentic tasks: ALFWorld embodied household tasks, BabyAI grid-world navigation, bundled web shopping, ScienceWorld interactive science, Jericho text-adventure games, and group travel planning. The dataset is paired with two evaluation suites: Reflection Evaluation and Evolution Evaluation. Each episode includes a unique identifier, environment/game instance, task type, agent model, success status, snapshot (trajectory context for reflection), and annotated failure instances with detection, localization, and diagnosis labels. The data is derived from several agentic environments under their respective licenses, with original annotations released under CC-BY-4.0.




