TOOLSANDBOX
收藏资源简介:
TOOLSANDBOX是由苹果公司开发的一个用于评估大型语言模型(LLM)工具使用能力的数据集。该数据集包含1032个精心设计的测试案例,涉及复杂的工具使用场景,如状态依赖、规范化及信息不足等。数据集创建过程中,引入了隐式状态依赖、LLM模拟用户和动态评估策略等创新元素。TOOLSANDBOX主要应用于评估和提升LLM在实际任务中的工具使用能力,特别是在需要复杂交互和状态管理的对话系统中。
TOOLSANDBOX is a dataset developed by Apple Inc. for evaluating the tool-use capabilities of Large Language Models (LLMs). This dataset contains 1,032 meticulously designed test cases covering complex tool-use scenarios such as state dependency, normalization, and insufficient information. During the dataset's creation, innovative elements including implicit state dependency, LLM-simulated users, and dynamic evaluation strategies were introduced. TOOLSANDBOX is primarily applied to evaluate and enhance the tool-use performance of LLMs in real-world tasks, especially in dialogue systems requiring complex interactions and state management.




