ClassEval
收藏资源简介:
ClassEval是由复旦大学开发的手工构建的数据集,用于评估大型语言模型在类级别代码生成上的表现。该数据集包含100个类级别的Python代码生成任务,涉及约500个人工小时的工作。ClassEval覆盖了实际软件开发的广泛主题,如管理系统、游戏开发等。每个任务都设计有高测试充分性的测试套件,以确保生成的代码的正确性。数据集旨在解决现有评估在简单代码生成场景上的局限性,提供一个更复杂、更接近实际开发环境的评估基准。
ClassEval is a manually constructed dataset developed by Fudan University for evaluating the performance of large language models (LLMs) on class-level code generation tasks. This dataset includes 100 class-level Python code generation tasks, which required approximately 500 person-hours of manual work. ClassEval covers a broad spectrum of topics in real-world software development, such as management systems and game development. Each task is paired with a high-test-adequacy test suite to validate the correctness of the generated code. The dataset aims to overcome the limitations of existing evaluations that only target simple code generation scenarios, thereby providing a more complex and realistic evaluation benchmark that closely resembles actual software development workflows.

- 1ClassEval: A Manually-Crafted Benchmark for Evaluating LLMs on Class-level Code Generation复旦大学 · 2023年



