RV-Bench
收藏资源简介:
RV-Bench是由香港理工大学等机构提出的一个用于评估大语言模型数学推理能力的基准数据集。该数据集基于MATH和LeetCode-Math两个数据源构建,包含900多个随机变量问题。通过随机化变量组合,RV-Bench能够有效评估模型在解决数学问题时的真实推理能力,避免了现有基准测试中可能存在的数据泄露问题。数据集的构建过程包括问题函数的初始化、求解和生成模块,确保了问题的多样性和难度与原问题一致。RV-Bench旨在解决当前LLMs在复杂数学推理任务中的性能评估问题,为模型提供了更真实的测试环境。
RV-Bench is a benchmark dataset proposed by institutions including the Hong Kong Polytechnic University for evaluating the mathematical reasoning capabilities of Large Language Models (LLMs). Constructed based on two data sources, MATH and LeetCode-Math, this dataset contains over 900 random variable problems. By randomizing variable combinations, RV-Bench can effectively assess the genuine reasoning abilities of models when solving mathematical problems, while avoiding potential data leakage issues present in existing benchmark tests. The construction pipeline of the dataset includes modules for problem function initialization, solution, and generation, ensuring that the diversity and difficulty level of the generated problems match those of the original source problems. RV-Bench is designed to address the performance evaluation challenges of current LLMs in complex mathematical reasoning tasks, providing a more authentic testing environment for these models.

- 1Benchmarking Large Language Models via Random Variables香港理工大学, 电子科技大学, 暨南大学, 西蒙弗雷泽大学 · 2025年



