遇见数据集

A leakage-controlled phishing benchmark with forensic structural features and frozen evaluation splits

收藏
Zenodo2026-08-07 更新2026-08-13 收录
官方服务:

资源简介:

A leakage-controlled benchmark of 5,444 live websites (2,569 phishing, 2,875 legitimate) for phishing detection, together with 78 forensic structural features that are invariant to URL rewriting, and SAIL, a selective adaptive incremental learner that gates adaptation on two label-free trust statistics. Phishing benchmarks are commonly assembled by pairing deep-linked URLs from a feed such as PhishTank against bare homepages from a ranking such as Tranco. A classifier can then separate the classes almost perfectly from the URL string alone, because it is learning the data source rather than maliciousness. Applying hostname-grouped splitting does not fix this: on a naively paired corpus every model tested still scored 0.940-0.996 F1 after grouping. This benchmark applies four construction controls so that the two classes are not separable by provenance. The deposit contains the extracted feature matrix, per-site metadata, the 242-domain redirect block-list, frozen train/validation/test split indices for both the hostname-grouped and chronological protocols, the full experimental pipeline, and a consistency checker that re-derives every number in the accompanying article from its generating artifact. Raw page content is not redistributed, as it is third-party material; source URLs are included so the corpus can be re-crawled."

提供机构:
Zenodo
创建时间:
2026-08-07
二维码
社区交流群
二维码
科研交流群
商业服务