HE-R, HE-R+, MBPP-R, MBPP-R+
收藏资源简介:
本文介绍了HE-R、HE-R+、MBPP-R和MBPP-R+四个数据集,这些数据集是由HumanEval和Mostly Basic Programming Problems (MBPP)改编而来,用于评估合成验证方法在评估解决方案正确性方面的影响。这些数据集将现有的编码基准测试转化为评分和排名数据集,以评估合成验证方法的有效性。数据集的具体大小、数据量等信息未在摘要中详细说明,但提到了这些数据集能够评估大型语言模型在代码测试用例生成方面的能力,并用于比较不同合成验证方法的性能。
This paper presents four datasets: HE-R, HE-R+, MBPP-R, and MBPP-R+, which are adapted from HumanEval and Mostly Basic Programming Problems (MBPP). These datasets are designed to evaluate the impact of synthetic validation methods on the assessment of solution correctness. These datasets transform existing coding benchmark tests into scoring and ranking datasets for evaluating the effectiveness of synthetic validation methods. The specific details such as the size and volume of the datasets are not elaborated in the abstract. Nevertheless, it is noted that these datasets can evaluate the ability of large language models (LLMs) to generate code test cases, and are utilized to compare the performance of different synthetic validation methods.

- 1Scoring Verifiers: Evaluating Synthetic Verification in Code and ReasoningNVIDIA Santa Clara, CA 15213, USA · 2025年



