LLM-Based Multi-Agent Systems for Code Generation: A Multi-Vocal Literature Review
收藏资源简介:
This dataset is associated with the paper titled "LLM-Based Agents for Code Generation: A Multi-Vocal Literature Review." The goal of this study is to systematically investigate the use of LLM-based agentic systems in code generation by analyzing both peer-reviewed and grey literature. The study aims to identify commonly used models, benchmarks, architectural designs, challenges, proposed solutions, and future research directions in this domain. The data is stored in a Microsoft Excel file consisting of several worksheets. Below is a simplified description of each worksheet. Summary This sheet provides an overall summary of both peer-reviewed and grey literature. It includes the initial results obtained from title and keyword searches, abstract screening, and full-paper review. Research Questions This sheet contains the research questions of the study along with the complete search strings used during the literature search process. Grey Literature Data This sheet contains all grey literature data retrieved from three sources: Google, Yahoo, and Bing. The initial search yielded 482 studies from these sources. Grey Literature Process This sheet documents the selection process for grey literature. After removing 84 duplicate studies from the initial 482 results, 398 studies remained. The sheet presents the filtering stages, including title and keyword screening and full-text review. Grey color represents studies selected for final full-paper reading. Green color represents studies selected after title and keyword screening but excluded during full-paper review.The sheet also includes quality assessment criteria, through which two papers were excluded.Additionally, three papers identified through the snowballing process are included at the end of this sheet. Grey Literature Data Extraction This sheet contains the detailed data extraction for the final 40 selected grey literature studies. Peer-Reviewed Studies Process This sheet describes the selection process for peer-reviewed papers collected from five digital libraries and through backward and forward snowballing. Grey-colored papers indicate inclusion for full-paper screening. Green-highlighted papers indicate inclusion after abstract screening but exclusion after full-paper review. Blue represents the paper included after titel and keywords reading. For search engines such as Google Scholar (which initially returned 918 results), only studies selected after title screening were stored, as Google Scholar does not allow full export of search results. Similarly, only relevant studies were added during backward and forward snowballing. Peer-Reviewed Data Extraction This sheet contains detailed data extraction for the finalized 74 peer-reviewed studies. Demographic Data Analysis This sheet provides demographic information about the selected studies, including authors’ affiliations, publication years, types of studies, and publication venues. Reasons Analysis This sheet presents the analysis results for RQ2, identifying the reasons for adopting agent-based systems in code generation. Model Analysis This sheet lists and analyzes all models utilized by the selected studies to implement agentic systems. Benchmarks Analysis This sheet provides the list and analysis of benchmarks used by the selected studies to evaluate agent-based code generation systems. Challenges and Solutions Analysis This sheet presents the identified challenges in agent-based code generation systems along with the proposed solutions reported in the literature. Future Work Analysis This sheet summarizes the future research directions identified in the selected studies.
本数据集关联于题为《基于大语言模型的智能体用于代码生成:一项多视角文献综述》(LLM-Based Agents for Code Generation: A Multi-Vocal Literature Review)的论文。本研究旨在通过分析同行评议文献与灰色文献,系统性探究基于大语言模型(LLM)的智能体系统在代码生成领域的应用。本研究目标为识别该领域内常用的模型、基准测试集、架构设计、现存挑战、已提出的解决方案以及未来研究方向。 本数据以Microsoft Excel文件形式存储,包含多个工作表。以下为各工作表的简化说明。 总览工作表(Summary):本工作表涵盖同行评议文献与灰色文献的整体概况,包含从标题与关键词检索、摘要筛选到全文审阅的全部初始结果。 研究问题工作表(Research Questions):本工作表收录了本研究的全部研究问题,以及文献检索过程中使用的完整检索式。 灰色文献数据表(Grey Literature Data):本工作表包含从Google、Yahoo及Bing三个检索源获取的全部灰色文献数据。本次初始检索从上述来源共得到482项研究。 灰色文献筛选流程工作表(Grey Literature Process):本工作表记录了灰色文献的筛选流程。初始482项结果中移除84项重复研究后,剩余398项。工作表展示了筛选阶段,包括标题与关键词筛选及全文审阅。 灰色标注代表入选最终全文阅读环节的研究。 绿色标注代表通过标题与关键词筛选后入选,但在全文审阅阶段被排除的研究。本工作表还包含质量评估标准,依据该标准共排除2篇文献。此外,通过滚雪球检索得到的3篇文献也被收录至本工作表末尾。 灰色文献数据提取工作表(Grey Literature Data Extraction):本工作表包含最终入选的40项灰色文献的详细数据提取内容。 同行评议文献筛选流程工作表(Peer-Reviewed Studies Process):本工作表描述了从5个数字图书馆以及通过正向、反向滚雪球检索收集到的同行评议论文的筛选流程。 灰色标注的论文代表入选全文筛选环节的文献。 绿色高亮的论文代表通过摘要筛选后入选,但在全文审阅阶段被排除的文献。 蓝色标注代表通过标题与关键词审阅后入选的文献。 对于Google Scholar等检索引擎(初始返回918项结果),由于其不支持完整导出检索结果,仅存储通过标题筛选后的研究。同理,在正向与反向滚雪球检索过程中,仅添加相关研究。 同行评议文献数据提取工作表(Peer-Reviewed Data Extraction):本工作表包含最终入选的74项同行评议文献的详细数据提取内容。 文献特征数据分析工作表(Demographic Data Analysis):本工作表提供入选文献的特征信息,包括作者所属机构、发表年份、研究类型及发表渠道。 采用动因分析工作表(Reasons Analysis):本工作表展示了针对研究问题2的分析结果,明确了在代码生成中采用智能体系统的动因。 模型分析工作表(Model Analysis):本工作表列出并分析了入选研究中用于构建智能体系统的全部模型。 基准测试集分析工作表(Benchmarks Analysis):本工作表列出并分析了入选研究中用于评估智能体代码生成系统的基准测试集。 挑战与解决方案分析工作表(Challenges and Solutions Analysis):本工作表梳理了智能体代码生成系统中已识别的挑战,以及文献中提出的对应解决方案。 未来研究方向分析工作表(Future Work Analysis):本工作表总结了入选研究中提出的未来研究方向。



