自动生成的代码任务基准
收藏资源简介:
自动生成的代码任务基准是由IBM以色列的研究团队开发的一个用于评估和验证代码相关任务解决方案的数据集。该数据集包含多种编程语言和代码任务的样本,如代码翻译、生成、完成、测试生成和摘要等。数据集的创建过程利用了图表示法和链式LLM代理,通过循环生成和验证代码相关工件。该数据集主要用于早期测试和验证LLM解决方案的有用性,旨在解决代码生成任务中的质量评估问题。
The automatically generated code task benchmark is a dataset developed by the research team at IBM Israel for evaluating and validating solutions to code-related tasks. This dataset includes samples across multiple programming languages and various code tasks, such as code translation, code generation, code completion, test generation, and code summarization. The dataset's development process leverages graph representation and chained LLM agents to iteratively generate and validate code-related artifacts. It is primarily used for early-stage testing and validation of the utility of LLM-based solutions, with the goal of addressing the quality evaluation issue in code generation tasks.

- 1Automatic Generation of Benchmarks and Reliable LLM Judgment for Code TasksIBM以色列 · 2024年



