eyeballvul
收藏资源简介:
eyeballvul数据集是由Timothée Chauvin创建,专门用于测试语言模型在实际环境中检测漏洞的能力。该数据集包含超过24,000个真实世界的安全漏洞,涵盖6,000多个修订版本和5,000多个开源仓库,总大小约为55GB。数据集的创建过程包括从OSV数据集中下载与开源仓库相关的CVEs,并通过一系列步骤将其转换为适合评估的格式。eyeballvul数据集主要应用于评估和提升语言模型在静态应用安全测试(SAST)中的性能,旨在解决大规模代码库中安全漏洞检测的问题。
The eyeballvul dataset was developed by Timothée Chauvin, specifically tailored to assess the vulnerability detection capabilities of language models in real-world scenarios. This dataset contains over 24,000 real-world security vulnerabilities, covering more than 6,000 code revisions and over 5,000 open-source repositories, with a total size of approximately 55 GB. The dataset was constructed by downloading CVEs associated with open-source repositories from the OSV dataset, then converting these CVEs into evaluation-ready formats through a series of standardized steps. The eyeballvul dataset is primarily used to evaluate and enhance the performance of language models in Static Application Security Testing (SAST), aiming to address the challenges of security vulnerability detection in large-scale codebases.




