遇见数据集

juxam/C3-VULMAP: C3-VULMAP Dataset v1.0

收藏
Zenodo2026-04-06 更新2026-05-26 收录
官方服务:

资源简介:

C3-VULMAP is a large-scale dataset developed to support research in Software Security, Vulnerability Detection, and Machine Learning for secure software development. The dataset contains millions of labeled source code samples mapped to standardized vulnerability categories derived from the Common Weakness Enumeration (CWE) taxonomy. The dataset includes 7,910,174 code samples annotated with vulnerability-related metadata and classification labels. Each entry contains a source code snippet together with its associated vulnerability identifier, vulnerability category, and additional descriptive attributes. The dataset is structured to enable both binary vulnerability detection and multi-class vulnerability classification tasks. C3-VULMAP was designed to facilitate research in automated vulnerability detection using both traditional static analysis techniques and modern machine learning approaches. In particular, the dataset supports experiments involving structural code representations derived from program analysis methods such as the Abstract Syntax Tree (AST) and graph-based learning techniques. The dataset contains 102,038 vulnerable samples (1.3%) and 7,808,136 non-vulnerable samples (98.7%), reflecting the class imbalance typically observed in real-world software vulnerability datasets. The dataset also includes 775 distinct CWE identifiers, enabling detailed studies of vulnerability types and categories across a large corpus of code. To support efficient large-scale data analysis and machine learning pipelines, C3-VULMAP is distributed in both CSV and Parquet formats. The Parquet version enables high-performance data processing for large-scale experiments, while the CSV format ensures accessibility for a wide range of data analysis tools. C3-VULMAP aims to facilitate reproducible research and benchmarking of automated vulnerability detection techniques by providing a large and diverse dataset for evaluating machine learning models and program analysis methods. Dataset Characteristics Property Value Total Samples 7,910,174 Vulnerable Samples 102,038 Non-Vulnerable Samples 7,808,136 Unique CWE IDs 775 Number of Attributes 9 Unique Code Samples 7,663,589 Dataset Formats Parquet (~16 GB)

C3-VULMAP是一款专为支撑软件安全、漏洞检测以及安全软件开发领域机器学习研究而构建的大规模数据集。该数据集包含数百万条标注源代码样本,这些样本映射自通用弱点枚举(Common Weakness Enumeration, CWE)分类体系下的标准化漏洞类别。 该数据集共包含7,910,174条附带漏洞相关元数据与分类标签的代码样本。每条数据条目均包含一段源代码片段,以及其对应的漏洞标识符、漏洞类别与其他描述性属性。本数据集的结构可同时支持二元漏洞检测与多分类漏洞分类两类任务。 C3-VULMAP的设计初衷是推动基于传统静态分析技术与现代机器学习方法的自动化漏洞检测研究,尤其支持依托程序分析方法(如抽象语法树(Abstract Syntax Tree, AST))与基于图的学习技术所生成的结构化代码表示的相关实验。 数据集中包含102,038条存在漏洞的样本(占比1.3%)与7,808,136条无漏洞样本(占比98.7%),这一分布反映了现实世界软件漏洞数据集中常见的类别不平衡问题。此外,该数据集涵盖775个不同的CWE标识符,可支撑针对大规模代码语料库中漏洞类型与类别的精细化研究。 为支持高效的大规模数据分析与机器学习流水线构建,C3-VULMAP以CSV与Parquet两种格式发布。其中Parquet格式可实现大规模实验的高性能数据处理,而CSV格式则可兼容各类主流数据分析工具,保障了数据集的普适易用性。 C3-VULMAP旨在通过提供大规模且多样化的数据集,用于评估机器学习模型与程序分析方法,推动自动化漏洞检测技术的可复现研究与基准测试工作。 数据集特征 属性 属性值 总样本数 7,910,174 漏洞样本数 102,038 无漏洞样本数 7,808,136 唯一CWE标识符数量 775 属性数量 9 唯一代码样本数 7,663,589 数据集格式 Parquet(约16 GB)

提供机构:
Zenodo
创建时间:
2026-04-06
二维码
社区交流群
二维码
科研交流群
商业服务