ToolBench-X
收藏资源简介:
ToolBench-X是由上海交通大学人工智能学院研发的综合性基准数据集,旨在系统评估智能体在不可靠工具环境下的鲁棒性。该数据集包含1106个可执行的多步骤任务,涵盖信息通信、商业交易、娱乐媒体等七大领域,并整合了顺序、并行及混合工作流,每个任务均配备确定性工具和标准答案以实现自动化评估。其构建过程始于清洁工具环境,随后通过注入规范漂移、调用错误、执行失败、输出漂移和跨源冲突等五类结构化可靠性风险,确保每个任务至少存在一条有效恢复路径。该数据集主要应用于人工智能领域,旨在解决现有评估基准对工具环境不可靠性考量不足的问题,推动智能体从单纯函数调用向复杂任务完成的可靠性评估演进。
ToolBench-X is a comprehensive benchmark dataset developed by the School of Artificial Intelligence at Shanghai Jiao Tong University, designed to systematically evaluate the robustness of AI Agents in unreliable tool environments. This dataset consists of 1,106 executable multi-step tasks spanning seven major domains such as information and communication, business transactions, and entertainment media. It integrates sequential, parallel, and hybrid workflows, with each task equipped with deterministic tools and standard answers to enable automated evaluation. Its construction starts from a clean tool environment, followed by injecting five types of structured reliability risks including specification drift, invocation errors, execution failures, output drift and cross-source conflicts, to ensure that each task has at least one valid recovery path. This dataset is primarily applied in the field of artificial intelligence, aiming to address the insufficient consideration of tool environment unreliability in existing evaluation benchmarks, and promote the advancement of AI Agent reliability assessment from simple function call-based evaluation to reliability assessment for complex task completion.





