ProBench
收藏资源简介:
ProBench是一个针对竞赛编程的大型语言模型评估基准,由天津大学提出。该数据集从Codeforces、Luogu和Nowcoder三个编程竞赛平台收集了2024年7月至12月期间的竞赛问题,通过在线提交生成的代码解决方案,利用原竞赛平台的全面测试用例来严格评估代码的正确性。数据集包含问题描述、难度等级和算法标签等详细信息,旨在全面、公平、深入地分析大型语言模型在竞赛编程中的推理能力。
ProBench is a large language model evaluation benchmark for competitive programming, proposed by Tianjin University. This dataset collects competitive programming problems from three platforms—Codeforces, Luogu, and Nowcoder—covering the period from July to December 2024. It uses code solutions generated via online submissions and rigorously evaluates the correctness of the code with the comprehensive test cases provided by the original contest platforms. The dataset contains detailed information including problem descriptions, difficulty levels, and algorithm tags, and aims to conduct a comprehensive, fair, and in-depth analysis of the reasoning capabilities of large language models in competitive programming.




