FeedbackEval
收藏资源简介:
FeedbackEval是一个用于评估大型语言模型在反馈驱动的代码修复任务中的性能的基准。该数据集由中山大学软件工程学院创建,包含了394个编码任务,涵盖了多种编程场景,每个任务都包含了错误代码段和四种不同类型的反馈。数据集旨在模拟真实软件开发场景中开发者从各种来源接收到的指导,包括结构化的测试反馈和编译器反馈,以及非结构化的人类反馈和简单反馈。
FeedbackEval is a benchmark for evaluating the performance of large language models (LLMs) on feedback-driven code repair tasks. This dataset was created by the School of Software Engineering, Sun Yat-sen University, and contains 394 coding tasks covering diverse programming scenarios. Each task includes erroneous code snippets and four distinct types of feedback. The dataset aims to simulate the guidance that developers receive from various sources in real-world software development scenarios, including structured test feedback, compiler feedback, as well as unstructured human feedback and simple feedback.

- 1FeedbackEval: A Benchmark for Evaluating Large Language Models in Feedback-Driven Code Repair Tasks中山大学软件工程学院 · 2025年



