Combi-Puzzles
收藏资源简介:
Combi-Puzzles数据集由基辅塔拉斯舍甫琴科国立大学和剑桥大学的研究人员创建,包含125个基于25个组合推理问题的变体,旨在评估大型语言模型(LLMs)在组合数学问题上的推理能力。数据集通过系统地操纵问题陈述,创建了五种不同形式的问题变体,包括数学形式、对抗性添加、参数化变化和语言混淆,以测试模型和人类在不同问题表述下的表现。数据集的创建过程严格控制,确保数学核心不变,适用于评估LLMs的泛化能力和人类在组合数学问题上的表现。
The Combi-Puzzles Dataset was created by researchers from Taras Shevchenko National University of Kyiv and the University of Cambridge. It comprises 125 variants based on 25 combinatorial reasoning problems, with the primary goal of evaluating the reasoning proficiency of Large Language Models (LLMs) on combinatorial mathematics problems. To test the performance of both models and humans under varied problem formulations, the dataset systematically manipulates problem statements to produce five distinct types of problem variants, including mathematical formalization, adversarial augmentation, parametric variation, and linguistic obfuscation. The dataset was developed under strict control to ensure the retention of the underlying mathematical core, making it a suitable benchmark for assessing both the generalization ability of LLMs and human performance on combinatorial mathematics tasks.

- 1Can Language Models Rival Mathematics Students? Evaluating Mathematical Reasoning through Textual Manipulation and Human Experiments基辅塔拉斯舍甫琴科国立大学, 剑桥大学 · 2024年



