CombiBench
收藏资源简介:
CombiBench是一个包含100个组合数学问题的综合基准测试集,每个问题都使用Lean 4进行了形式化,并配有相应的非正式描述。问题集涵盖了从初中到IMO和大学水平的各种难度,涵盖了十多个组合数学主题。CombiBench旨在解决自动定理证明领域中组合数学缺乏适当基准和定理库的问题。该数据集包含了自2000年以来所有IMO组合数学问题(除了2004年的第3题,因为其陈述包含图像)。此外,我们提供了一个全面的标准评估框架,称为Fine-Eval,用于形式数学的评估。它不仅适用于基于证明的问题,而且首次支持对填空题的评估。使用Fine-Eval作为评估方法,并以Kimina Lean Server为后端,我们在CombiBench上对几个LLM进行了基准测试,并观察到它们在形式化解决组合数学问题方面的能力仍然有限。在所有测试的模型中(没有一个是为这个特定任务训练的),Kimina-Prover取得了最佳结果,在“有解决方案”和“无解决方案”的情况下都解决了7个问题(共100个)。我们开源了基准数据集以及所提出的评估方法的代码。
CombiBench is a comprehensive benchmark dataset consisting of 100 combinatorial mathematics problems. Each problem has been formalized using Lean 4, paired with corresponding informal descriptions. The problem set spans a wide range of difficulty levels, from middle school mathematics, through International Mathematical Olympiad (IMO) and up to university-level content, covering more than ten combinatorial mathematics topics. CombiBench is designed to address the gap in appropriate benchmarks and theorem libraries for combinatorial mathematics within the field of automated theorem proving. This dataset includes all IMO combinatorial problems since 2000, with the exception of Problem 3 from the 2004 IMO, as its problem statement contains an image. Additionally, we present a comprehensive standard evaluation framework named Fine-Eval for formal mathematics. This framework is not only applicable to proof-based problems, but also supports the evaluation of fill-in-the-blank questions for the first time. Using Fine-Eval as the evaluation methodology and Kimina Lean Server as the backend, we benchmarked several large language models (LLMs) on CombiBench, and observed that their capabilities in formally solving combinatorial mathematics problems remain limited. Among all tested models—none of which were trained specifically for this task—Kimina-Prover achieved the best performance, solving 7 out of 100 problems across both the "with solution" and "without solution" evaluation settings. We have open-sourced both the benchmark dataset and the code for the proposed evaluation framework.

- 1CombiBench: Benchmarking LLM Capability for Combinatorial Mathematics中国科学院数学与系统科学研究院, 中山大学, 剑桥大学, 华东师范大学, 伦敦帝国学院, 斯德哥尔摩大学, Numina, 月之船人工智能公司 · 2025年



