C3-VULMAP: A Large-Scale Dataset for Software Vulnerability Detection (v3.0 — Provenance-Rebuilt Two-Tier Corpus)
收藏资源简介:
C3-VULMAP is a large-scale dataset developed to support research in Software Security, Vulnerability Detection, and Machine Learning for secure software development. The dataset maps C/C++ source code to standardized vulnerability categories derived from the Common Weakness Enumeration (CWE) taxonomy, further mapped to LINDDUN privacy threat types — so findings carry an explicit privacy dimension, not just a CWE tag. This version is a substantial rebuild of the original release described in the companion Electronics paper (doi: 10.3390/electronics14132703). The single pooled file from that release has been replaced with an explicitly two-tier structure, and full source-level provenance — discarded in the original construction — is now retained on every row. See "What changed" below. C3-VULMAP was designed to facilitate research in automated vulnerability detection using both traditional static analysis techniques and modern machine learning approaches, including structural code representations such as Abstract Syntax Trees (AST) and graph-based learning. Two tiers, released separately, never pooled Tier 1 — Supervised corpus (corpus_final.parquet): every label is backed by a real, published CVE or commit record. Tier 2 — Static-analysis labels (starcoder_tier2.parquet + starcoder_tier2_static_labels.csv): general C/C++ source with heuristic CWE pattern-matches attached by static analysis, not by a verified report. A file with no flagged pattern in this tier is not thereby confirmed safe — absence of a match is not evidence of absence of a vulnerability, and this tier should never be treated as a verified-negative class. Every row states which kind of claim it carries via a label_source column (cve_verified vs. static_analysis). Dataset Characteristics Property Tier 1 (supervised) Tier 2 (static-analysis) Total rows 332,970 14,838,026 Vulnerable (CVE/commit-verified) 18,365 (5.5%) n/a — no verified label Non-vulnerable (CVE/commit-verified safe) 314,605 (94.5%) n/a — no verified label Files with a static-analysis-flagged pattern n/a 2,102,354 Unique CWE identifiers 440 440 (same mapping chain) Number of attributes 10 8 (labels file) Dataset formats Parquet Parquet (source) + CSV (labels) Files corpus_final.parquet — Tier 1, 332,970 rows starcoder_tier2.parquet — Tier 2 source text, 14,838,026 rows starcoder_tier2_static_labels.csv — Tier 2 static-analysis labels, 2,102,354 flagged files cwe_category_map.csv, category_id_to_linddun.csv, excluded_abstract_cwes.csv — the full CWE→category→LINDDUN mapping chain, for independent verification What changed from the original (v1) release Replaced one pooled file with an explicitly separated two-tier structure and a label_source column stating the evidentiary basis of every label, so the two can never be silently conflated. Retained full source-level provenance (source dataset, project, commit, file path) on every row — the original release discarded this after construction. Added 140 targeted synthetic examples, written to fill the LINDDUN categories the real data left thinnest, verified against the same contamination screen as the rest of the corpus. Corrected the CWE→LINDDUN mapping chain: excluded 65 CWEs MITRE itself designates too abstract to map to a specific vulnerability (checked directly against MITRE's own definitions), added coverage MITRE explicitly endorses, and added five newly-curated category-level LINDDUN mappings. Added an independent, clearly-separated 14.8M-row static-analysis tier, rather than pooling unverified data into the supervised corpus — the specific defect this rebuild exists to correct. Full construction methodology, with every number independently reproducible from the source scripts: https://github.com/juxam/C3-VULMAP/blob/main/CONSTRUCTION.md



