遇见数据集

Verification data for "How You Compose Matters: Necessary Conditions for Corroborated Promotion and False-Positive Structure in Operational Webshell Detection"

收藏
Zenodo2026-08-17 更新2026-08-20 收录
官方服务:

资源简介:

Superseded by Version 3.0.0 (10.5281/zenodo.21980688) in the same record. This version was published during upload before the 3.0.0 bundle was attached; its only file is carried over from version 2.0.0. Use Version 3.0.0. Per-sample prediction tables and recomputation scripts supporting every table reported in the accompanying paper. The archive contains the adjudication record for each file in the evaluation corpora — detector score, verification-operator verdict, final verdict after the eighteen post-processing stages, the three clauses of the Bounded Promotion invariant, invariant violation flags, structural evidence tags, and a per-stage verdict snapshot — together with the scripts that recompute the published tables from those records. Five measurement conditions are included, each over five random seeds: the deployed configuration, the same system with the promotion admissibility test disabled, the same benchmark under the previous training corpus, a flattened benchmark in which all path information is removed, and a post-split corpus of webshells whose first commit postdates the training freeze. The corpus construction criteria and the engine patch that installs the admissibility gate are included, so that each reported ablation can be reproduced. All tables in the paper can be recomputed from these files without access to the detection model or the detection engine. Neither is included: the model weights and the engine are commercial artifacts. Webshell samples are not redistributed in accordance with source platform policy. Provenance coverage differs by corpus. The evaluation benchmark (1,901 files) and the post-split corpus (98 files) carry a SHA-256 hash and source repository attribution for every file, so both can be reconstructed from public repositories and verified by hash; the post-split manifest additionally records each file's git first-commit date. The training corpus manifest (128,297 files) provides SHA-256, label and language for every file, but source repository attribution only for the most recent augmentation — the remainder were renamed to content hashes before the manifest was assembled. All results reported in the paper are computed on the evaluation corpora. Supersedes v1.0.0, which corresponds to an earlier benchmark construction not reported in the published article.

提供机构:
Zenodo
创建时间:
2026-08-17
二维码
社区交流群
二维码
科研交流群
商业服务