MATH-Perturb
收藏资源简介:
MATH-Perturb数据集是由普林斯顿大学和谷歌的研究人员构建的,包含279个经过简单扰动和困难扰动的数学问题,这些问题源自MATH数据集的第五难度级别(最难的问题)。数据集通过12位具有强大数学背景的博士生进行注释和交叉验证,确保了质量。该数据集旨在评估大型语言模型在数学推理任务中的鲁棒性,特别是面对问题公式的基本变化时的表现。
The MATH-Perturb dataset was constructed by researchers from Princeton University and Google. It contains 279 math problems with both simple and difficult perturbations, which are derived from the 5th difficulty level (the hardest problems) of the MATH dataset. The dataset was annotated and cross-validated by 12 doctoral candidates with strong mathematical backgrounds to ensure its quality. This dataset aims to evaluate the robustness of large language models (LLMs) in mathematical reasoning tasks, particularly their performance when facing fundamental changes to problem formulations.

- 1MATH-Perturb: Benchmarking LLMs' Math Reasoning Abilities against Hard Perturbations普林斯顿大学,谷歌 · 2025年



