NumericBench
收藏资源简介:
NumericBench是一个全面评估大型语言模型数值推理能力的基准,由香港理工大学、香港科技大学和香港中文大学的研究人员提出。该数据集涵盖了从合成数字列表到爬取的实时世界数据,旨在评估LLM在处理长期上下文、噪声和多步骤推理方面的挑战。它综合了六个数据集,包括算术数字、混合数字字符串、数字列表、股票、天气和数值序列模式,以评估LLM在数值识别、算术运算、上下文检索、比较、汇总和逻辑推理六个基本的数值推理能力。
NumericBench is a benchmark designed for comprehensive evaluation of the numerical reasoning capabilities of large language models (LLMs), proposed by researchers from The Hong Kong Polytechnic University, The Hong Kong University of Science and Technology, and The Chinese University of Hong Kong. This dataset covers content ranging from synthetic numerical lists to crawled real-time real-world data, aiming to evaluate the challenges that LLMs face when processing long contexts, noisy data, and multi-step reasoning tasks. It integrates six constituent datasets, including arithmetic numerals, mixed numerical strings, numerical lists, stock data, weather data, and numerical sequence patterns, to assess six core numerical reasoning abilities of LLMs: numerical recognition, arithmetic operations, contextual retrieval, comparison, summarization, and logical reasoning.
数据集概述
数据集名称
NumericBench
数据集简介
NumericBench是一个全面性的基准测试,旨在评估大型语言模型(LLM)的数值推理能力。该数据集涵盖了从算术运算到数字识别、上下文检索、比较、总结和逻辑推理等任务,以解决LLM在数值处理方面的局限性。数据集包含了从合成数字列表到股票趋势、天气模式等现实世界领域的多样化数据集,系统地在结构化和嘈杂的环境中对LLM进行测试。
数据集构成
- 包含多样化数据集,涵盖合成数字列表和现实世界领域数据。
- 用于评估LLM在数值推理方面的性能。
实验结果
- 实验结果显示,如GPT-4o和DeepSeek-V3等模型在数值推理方面存在显著弱点。
引用信息
@misc{li2025exposingnumeracygapsbenchmark, title={Exposing Numeracy Gaps: A Benchmark to Evaluate Fundamental Numerical Abilities in Large Language Models}, author={Haoyang Li and Xuejia Chen and Zhanchao XU and Darian Li and Nicole Hu and Fei Teng and Yiming Li and Luyu Qiu and Chen Jason Zhang and Qing Li and Lei Chen}, year={2025}, eprint={2502.11075}, archivePrefix={arXiv}, primaryClass={cs.CL}, url={https://arxiv.org/abs/2502.11075}, }




