MCP-RADAR
收藏资源简介:
MCP-RADAR是一个用于评估大型语言模型在模型上下文协议(MCP)框架下工具使用能力的全面基准数据集。该数据集涵盖了软件工程、数学推理和一般问题解决等三个核心领域,共包含300个任务。数据集旨在通过五个维度来衡量模型能力:答案准确性、工具选择效率、计算资源效率、参数构造准确性和执行速度。MCP-RADAR的构建过程经过严格的任务设计和验证,确保了数据集的高质量和可靠性。
MCP-RADAR is a comprehensive benchmark dataset for evaluating the tool-use capabilities of large language models under the Model Context Protocol (MCP) framework. This dataset covers three core domains including software engineering, mathematical reasoning, and general problem-solving, with a total of 300 tasks. It aims to evaluate model capabilities across five dimensions: answer accuracy, tool selection efficiency, computational resource efficiency, parameter construction accuracy, and execution speed. The construction process of MCP-RADAR has undergone rigorous task design and validation to ensure the high quality and reliability of the dataset.

- 1MCP-RADAR: A Multi-Dimensional Benchmark for Evaluating Tool Use Capabilities in Large Language Models西安交通大学 · 2025年



