tencent/C3-BenchMark
收藏资源简介:
C3-Bench是一个用于评估基于大型语言模型(LLM)的智能体在多任务处理中的鲁棒性的开源和高质量基准。它通过设计三个挑战(导航复杂的工具关系、处理关键隐藏信息和动态决策路径管理)来揭示模型在处理工具依赖、长上下文信息依赖和频繁策略类型切换方面的不足。C3-Bench还引入了细粒度的指标、创新的数据收集算法和可重复的评价方法。在49个主流智能体上的广泛实验表明,智能体在这些方面存在显著缺陷。C3-Bench旨在通过这些挑战暴露模型的漏洞,并推动对智能体性能可解释性的研究。
C3-Bench is an open-source and high-quality benchmark for evaluating the robustness of agents based on large language models (LLMs) in multitasking. It reveals the shortcomings of models in handling tool dependencies, long context information dependencies, and frequent policy-type switching by designing three challenges: navigating complex tool relationships, handling critical hidden information, and managing dynamic decision paths. C3-Bench introduces fine-grained metrics, innovative data collection algorithms, and reproducible evaluation methods. Extensive experiments on 49 mainstream agents show significant deficiencies in these aspects. C3-Bench aims to expose model vulnerabilities through these challenges and drive research into the interpretability of agent performance.




