UGMathBench
收藏资源简介:
UGMathBench是由香港科技大学数学系创建的一个多样化和动态的基准测试集,旨在评估大型语言模型(LLMs)在本科水平数学推理中的表现。该数据集包含5062个问题,涵盖16个学科和111个主题,具有10种不同的答案类型。每个问题包含三个随机化版本,以评估模型的推理鲁棒性。数据集来源于该机构的在线作业评分系统,经过数据收集、清理和去重等步骤生成。UGMathBench的应用领域主要是评估和改进LLMs在解决复杂数学问题中的推理能力,旨在解决现有基准测试集在覆盖范围和动态性方面的不足。
UGMathBench is a diverse and dynamic benchmark dataset created by the Department of Mathematics of The Hong Kong University of Science and Technology, designed to evaluate the performance of Large Language Models (LLMs) in undergraduate-level mathematical reasoning. This dataset contains 5,062 questions covering 16 disciplines and 111 topics, with 10 distinct answer types. Each question includes three randomized versions to assess the reasoning robustness of models. The dataset is derived from the institution's online homework grading system, and generated through processes including data collection, cleaning and deduplication. The main application fields of UGMathBench are evaluating and improving the reasoning abilities of LLMs when solving complex mathematical problems, aiming to address the shortcomings of existing benchmark datasets in terms of coverage and dynamism.

- 1UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language Models香港科技大学数学系 · 2025年



