NESTFUL
收藏资源简介:
NESTFUL数据集由IBM研究院创建,旨在评估大型语言模型(LLMs)在嵌套API调用序列中的能力。该数据集包含300个高质量的人工标注样本,分为可执行和不可执行两类。可执行样本通过爬取Rapid-APIs手动筛选,而不可执行样本则由人工从使用LLM生成的合成数据中挑选。数据集的创建过程强调了API调用的嵌套序列,旨在解决复杂的多步骤任务,特别是在需要多个API协同工作的实际应用场景中。
The NESTFUL dataset was developed by IBM Research to evaluate the capabilities of Large Language Models (LLMs) in nested API call sequences. This dataset comprises 300 high-quality manually annotated samples, categorized into two groups: executable and non-executable. Executable samples are manually curated by crawling Rapid-APIs, while non-executable samples are manually selected from synthetic data generated by LLMs. The development of the dataset prioritizes nested API call sequences, aiming to address complex multi-step tasks, especially in real-world application scenarios where multiple APIs need to work in tandem.

- 1NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API CallsIBM研究院 · 2024年



