Localization Is a Central Bottleneck for Small Local LLMs: A Gated Empirical Study of Vulnerability Remediation on Real Python CVEs
收藏资源简介:
This deposit contains the evaluation artifacts for "[Localization Is a Central Bottleneck for Small Local LLMs: A Gated Empirical Study of Vulnerability Remediation on Real Python CVEs]" — an exploratory, gated empirical study measuring whether five locally-hosted 2.7B–7B LLMs can detect, localize, and repair real vulnerabilities in 29 already-patched Python CVE instances (19 projects), without being given the vulnerable location in advance. Primary result: 9/145 gated trials (6.2%) correctly detected the vulnerable file; 2/145 (1.4%) correctly localized the vulnerable function; no 7B-class patch applied, compiled, and ran. Scale checks with Qwen3-Coder-30B, GPT-5.6, and Claude Pro are reported separately. A paired clean-control experiment (vulnerable vs. patched files) shows no discrimination signal above chance (MCC = −0.027). A security-regression-test oracle could not adjudicate any of the 83 applicable patches (0/19 candidate oracles currently validate); this reflects a coverage gap, not a demonstrated security failure. This version (updated [التاريخ]) adds the corrected oracle validator and its pre-fix backup, the dataset-correction verification and full rerun artifacts, and synchronizes the archive with the manuscript's V26_1 revision. Corresponding authors: dr.a.tengy@alexu.edu.eg; dr.ibrahim.6914r.adc@alexu.edu.eg



