PhishVN: A Time-Stamped Vietnamese URL Phishing Dataset with Impersonation-Scenario Labels and Confidence Tiers
收藏资源简介:
PhishVN is an open, time-stamped phishing-website dataset localised to the Vietnamese context. Its verified core is a URL table of 17,079 records (2,588 phishing, 14,491 legitimate), extended by an explicitly-tagged bronze expansion (34,283 community/feed-sourced phishing) for a full 51,362 records. Each record is described by a 21-feature lexical/infrastructure schema aligned with the CompPhish dataset to enable cross-dataset study. Phishing URLs are drawn from the National Cyber Security Centre feed (Tin Nhiem Mang; verified, gold/silver) together with the ChongLuaDao community blacklist and the OpenPhish feed (bronze); legitimate URLs combine the certified "trusted organisation" registry (curated, easy negatives), a .vn-filtered Tranco slice (popular Vietnamese-visited sites), and a global Tranco sample (harder, traffic-based negatives) to strengthen external validity and avoid the trivial "trusted-vs-malicious" and "non-.vn = phishing" shortcuts. Every record carries an impersonation-scenario label (bank, government, tax, e-commerce, telecom, delivery, social, gaming) inferred from brand tokens in the URL, and a label-confidence tier (gold = source-verified/handled; silver = under processing; bronze = undated community/feed expansion, excluded from the primary benchmark). Records are ordered by first-seen date and partitioned into train/validation/test using a group-aware temporal split keyed on the registrable domain, so near-duplicate campaign subdomains never span splits. Files: dataset_url.csv (full record table); vn_compphish.csv (same URLs re-featurised into the exact CompPhish schema); splits/url_{train,val,test}.csv; and docs/ (datasheet, column schema, data-source notes). A MANIFEST provides SHA-256 checksums. Notes: the is_https column is retained for schema compatibility but is a collection artefact and should be excluded from modelling (all lexical features are computed on the scheme-stripped URL). Personal data is redacted; no id-to-PII mapping is included. A preliminary multi-modal subset pairing captured page HTML (n=933) and screenshots (n=869) across both classes (209 phishing, 660 benign) is available separately as a research-only, gated tier. This dataset accompanies a companion cross-dataset detection study (in preparation). Please credit the upstream sources (NCSC Tin Nhiem Mang; ChongLuaDao; the Tranco list).




