Measuring the Wrong Thing: A Systematic Review of Evaluation Validity in Machine Learning for Vulnerability Management
收藏资源简介:
Replication package for the systematic review "Measuring the Wrong Thing? A Systematic Review of Evaluation Validity in Machine Learning for Vulnerability Management". The package reproduces every table, figure, and headline number in the paper from the shipped derived data. It contains the review protocol and research log; the extraction dataset (N = 379 studies, 2015–2026); all independent codings of the RQ3 both-reporter audit (an automated first pass, the author's blind full-text re-coding, and an independent second coder's screening, extraction, and written adjudication samples); the rank-test (RQ3.2) data; the round-2 search re-execution, recall audit, and scope-boundary artifacts; all LLM prompts and model identities; and stdlib-Python scripts with a one-command reproduce.sh plus a machine-verified number audit (verify_numbers.py). Copyright-reproduced full text of the reviewed papers is intentionally excluded; the analysis is reproducible from the shipped derived codings. See ZENODO_MANIFEST.md and README_REPRODUCE.md for the file map and instructions.



