RoTBench
收藏资源简介:
RoTBench是由复旦大学开发的一个多级基准数据集,用于评估大型语言模型在工具学习中的鲁棒性。该数据集包含五个不同噪音水平的环境,涵盖清洁、轻微、中等、重度和联合级别,旨在深入分析模型在工具选择、参数识别和内容填充三个关键阶段的稳定性。数据集通过模拟真实世界中的噪音情况,帮助研究人员评估和提高语言模型在复杂环境中的性能。
RoTBench is a multi-level benchmark dataset developed by Fudan University for evaluating the robustness of large language models (LLMs) in tool learning. This dataset includes five environments with distinct noise levels, covering clean, mild, moderate, severe, and combined noise scenarios, aiming to conduct in-depth analyses of model stability across three critical stages: tool selection, parameter identification, and content filling. By simulating real-world noise conditions, this dataset helps researchers evaluate and improve the performance of language models in complex environments.

- 1RoTBench: A Multi-Level Benchmark for Evaluating the Robustness of Large Language Models in Tool Learning复旦大学 · 2024年



