DeBinVul
收藏资源简介:
DeBinVul数据集由德克萨斯大学圣安东尼奥分校的Secure AI and Autonomy Laboratory创建,专注于C/C++语言,因其广泛应用于关键基础设施和与众多漏洞相关。该数据集包含150,872个样本,涵盖了多种CPU架构和编译优化级别,旨在识别、分类、描述和恢复反编译二进制代码中的漏洞。数据集的创建过程包括从多个来源收集源代码,进行编译和反编译,并使用GPT-4生成代码描述。DeBinVul数据集主要应用于增强大型语言模型在反编译二进制代码漏洞分析中的能力,旨在解决源代码与反编译二进制代码之间的语义差距问题。
The DeBinVul dataset was created by the Secure AI and Autonomy Laboratory at the University of Texas at San Antonio, focusing on the C/C++ programming languages, which are widely used in critical infrastructure and associated with numerous vulnerabilities. This dataset contains 150,872 samples covering multiple CPU architectures and compilation optimization levels, and is designed to identify, classify, characterize and recover vulnerabilities in decompiled binary code. The dataset creation process includes collecting source code from multiple sources, performing compilation and decompilation operations, and generating code descriptions via GPT-4. The DeBinVul dataset is primarily applied to enhance the capabilities of large language models (LLMs) in vulnerability analysis of decompiled binary code, aiming to address the semantic gap between source code and decompiled binary code.

- 1Enhancing Reverse Engineering: Investigating and Benchmarking Large Language Models for Vulnerability Analysis in Decompiled Binaries德克萨斯大学圣安东尼奥分校 · 2024年



