API-BLEND
收藏资源简介:
API-BLEND是一个为训练和系统测试工具增强型大型语言模型(LLMs)设计的大型语料库。该数据集模仿了涉及API任务的现实世界场景,如API/工具检测、槽填充和检测到的API的排序。API-BLEND由10个数据集组成,其中5个用于训练,5个用于域外测试,涵盖了语义解析、对话和数字助手等多个领域。通过混合方法生成数据,API-BLEND旨在解决现有数据集在API任务数据稀缺的问题,特别是序列化任务,并展示出比其他现有方法更好的域外泛化性能。
API-BLEND is a large-scale corpus designed for training and systematically testing tool-augmented Large Language Models (LLMs). It recreates real-world scenarios involving API-related tasks, such as API/tool detection, slot filling, and ranking of detected APIs. Comprising 10 individual datasets, API-BLEND includes 5 datasets for training and 5 for out-of-domain testing, covering multiple domains including semantic parsing, dialogue systems, and digital assistants. Using a hybrid data generation approach, API-BLEND aims to address the data scarcity issue of existing datasets for API-related tasks, particularly sequential tasks, and demonstrates better out-of-domain generalization performance compared to other existing methods.

- 1API-BLEND: A Comprehensive Corpora for Training and Benchmarking API LLMsIBM研究院 · 2024年



