MTU-Bench
收藏资源简介:
MTU-Bench是由阿里巴巴集团、中国科学院大学和滑铁卢大学联合创建的多粒度工具使用基准数据集,旨在评估大型语言模型在工具使用方面的能力。该数据集包含159,061条对话,涵盖了单轮单工具、单轮多工具、多轮单工具、多轮多工具以及分布外任务等多种场景。数据集通过转换现有高质量数据集来模拟真实世界的工具使用场景,并提出了一个名为MTU-Instruct的指令数据集以增强现有LLMs的工具使用能力。MTU-Bench的创建过程包括数据收集、工具创建、工具聚类、工具文档生成和工具使用数据合成等多个步骤。该数据集主要应用于提升大型语言模型在实际应用中的工具使用能力和解决复杂的工具调用问题。
MTU-Bench is a multi-granularity tool use benchmark dataset jointly created by Alibaba Group, University of Chinese Academy of Sciences, and University of Waterloo, aiming to evaluate the tool-use capabilities of large language models (LLMs). This dataset contains 159,061 dialogues, covering various scenarios including single-turn single-tool, single-turn multi-tool, multi-turn single-tool, multi-turn multi-tool, and out-of-distribution (OOD) tasks. By converting existing high-quality datasets, it simulates real-world tool use scenarios, and proposes an instruction dataset named MTU-Instruct to enhance the tool-use capabilities of existing LLMs. The development process of MTU-Bench comprises multiple steps such as data collection, tool creation, tool clustering, tool documentation generation, and tool use data synthesis. This dataset is primarily applied to improve the tool-use abilities of large language models in real-world applications and solve complex tool invocation problems.




