StableToolBench
收藏资源简介:
该数据集名为StableToolBench,是从ToolBench衍生出的评估数据集,旨在评估大型语言模型(LLMs)的工具学习能力。为了解决实时API可能带来的不稳定问题,该数据集还额外引入了缓存系统和API模拟器。此外,该数据集引入了两个关键指标:可解决通过率(SoPR)和可解决胜率(SoWR),以评估LLMs的表现。根据工具类别和场景,该数据集分为六个评估子集,用于大型语言模型的工具学习和评估任务。
This dataset, named StableToolBench, is an evaluation dataset derived from ToolBench, which is designed to assess the tool learning capabilities of large language models (LLMs). To mitigate the potential instability issues of real-time Application Programming Interfaces (APIs), this dataset additionally incorporates a caching system and API simulators. Furthermore, two key metrics, Solvable Pass Rate (SoPR) and Solvable Win Rate (SoWR), are introduced to evaluate the performance of LLMs. Based on tool categories and application scenarios, this dataset is split into six evaluation subsets for tool learning and evaluation tasks of large language models.

- 1StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models清华大学人工智能产业研究院 · 2024年



