HumanEval
收藏资源简介:
HumanEval是由代尔夫特理工大学的研究团队创建的软件工程领域AI基准数据集。该数据集主要用于评估大型语言模型在代码生成任务上的性能,包含多个子任务,如代码修复和代码解释。尽管有多个版本尝试扩展语言支持或改善测试覆盖范围,但原始数据集中的错误和不足仍未得到彻底解决。该数据集在AI和软件工程社区中广受欢迎,但存在测试不正确、测试覆盖不足、解决方案错误和问题描述不明确等问题。
HumanEval is an AI benchmark dataset in the field of software engineering, created by a research team from Delft University of Technology. It is primarily used to evaluate the performance of large language models on code generation tasks, and includes multiple subtasks such as code repair and code explanation. Although multiple revised versions have attempted to expand language support or improve test coverage, the errors and shortcomings in the original dataset have not been completely resolved. This dataset is widely popular in the AI and software engineering communities, but it still has issues including incorrect test cases, insufficient test coverage, erroneous solutions and ambiguous problem descriptions.




