Verification data for "How You Compose Matters: Necessary Conditions for Corroborated Promotion and False-Positive Structure in Operational Webshell Detection"
收藏资源简介:
Which file to use. This version contains two archives. Use zenodo_v2.zip — despite its filename it is the version 3.0.0 bundle. webshell-verification-data-v2.tar.gz is carried over from version 2.0.0 and is superseded. The README.md inside zenodo_v2.zip documents the bundle and the relation between versions. Per-sample prediction tables and recomputation scripts supporting every table reported in the accompanying paper. The archive contains the adjudication record for each file in the evaluation corpora — detector score, verification-operator verdict, final verdict after the eighteen post-processing stages, the three clauses of the Bounded Promotion invariant, invariant violation flags, structural evidence tags, and a per-stage verdict snapshot — together with the scripts that recompute the published tables from those records. Five measurement conditions are included, each over five random seeds: the deployed configuration, the same system with the promotion admissibility test disabled, the same benchmark under the previous training corpus, a flattened benchmark in which all path information is removed, and a post-split corpus of webshells whose first commit postdates the training freeze. Each run ships its verdict table together with the run environment and the invariant counts, so that the two engine builds compared in the admissibility ablation can be told apart by hash. The corpus construction criteria and the engine patch that installs the admissibility gate are included, so that each reported ablation can be reproduced. All tables in the paper can be recomputed from these files without access to the detection model or the detection engine. scripts/recompute_tables.py does so directly. Neither the model weights nor the engine is included: both are commercial artifacts. Webshell samples are not redistributed in accordance with source platform policy. Two conventions have to match before a difference in the fourth decimal means anything, and both are documented in the README: every rate in the paper is computed per seed and then averaged over the five seeds rather than from pooled counts, and the invariant checker evaluates the structural-evidence clause against a tag-level set rather than the disjunct-level definition. Both definitions are included so that either reading can be recomputed. Provenance coverage differs by corpus. The evaluation benchmark (1,901 files) and the post-split corpus (98 files) carry a SHA-256 hash and source repository attribution for every file, so both can be reconstructed from public repositories and verified by hash; the post-split manifest additionally records each file's git first-commit date. The training corpus manifest (128,297 files) provides SHA-256, label and language for every file, but source repository attribution only for the most recent augmentation — the remainder were renamed to content hashes before the manifest was assembled. All results reported in the paper are computed on the evaluation corpora. Two figures quoted in the paper predate this benchmark and cannot be recomputed from these files: the false-positive layer decomposition and the preprocessing-alignment counts were measured on an earlier development benchmark of 1,968 files. The experiment logs that produced them are deposited in logs/ in their place, and the README states which figures they are. Relation to earlier versions. Version 3.0.0 corresponds to the final article. Version 2.0.0 holds the same per-sample data but predates a verification pass that corrected several statements about it — the admissibility gate's selection rule, the canary count of the preprocessing parity test, and the aggregation convention behind the reported rates — and ships no recomputation script. Where the two disagree in wording, 3.0.0 is authoritative. Version 1.0.0 corresponds to an earlier benchmark construction that is not the evaluation corpus of the published article and is retained only for provenance.



