Synthetic Dataset for Code Vulnerability Flaws
收藏资源简介:
本研究提出了一个合成数据集,用于代码漏洞缺陷的代码审查。该数据集由Pontificia Universidad Católica de Chile的研究团队创建,旨在通过利用大型语言模型(LLMs)生成类似人类的代码审查评论,以解决现有数据集中安全相关审查样本不足的问题。数据集的内容基于安全漏洞相关的提交,包括提交的差异和相应的提交消息。研究团队计划使用这个合成数据集来微调现有的代码审查模型,并预期这将提高模型的性能。
This study proposes a synthetic dataset for code review focused on code vulnerability defects. Developed by the research team at Pontificia Universidad Católica de Chile, this dataset aims to address the shortage of security-related review samples in existing datasets by leveraging large language models (LLMs) to generate human-like code review comments. The dataset's content is based on security vulnerability-related commits, including commit diffs and their corresponding commit messages. The research team plans to use this synthetic dataset to fine-tune existing code review models, and expects that this will improve the models' performance.




