SynthIPv6-Diag: A Controlled, RFC-Informed Synthetic IPv6 Network-Flow Corpus for IDS Benchmarking and Membership-Inference Auditing
收藏资源简介:
SynthIPv6-Diag is a controlled, RFC-informed synthetic IPv6 network-flow corpus for reproducible intrusion-detection benchmarking and membership-inference auditing. Flows from three public IPv4 IDS corpora — CIC-IDS-2017, CIC-IDS-2018 and CIC-DDoS-2019 — are mapped to IPv6 address and transition templates (6to4, NAT64, Teredo and modified EUI-64, treated as distinct addressing/transition templates rather than interchangeable translators) and enriched with 27 IPv6-specific header fields synthesised from RFC-informed constraints, with a protocol-constraint manifest applied and minority classes rebalanced per class using SMOTE and a per-class WGAN-GP. The headline SynthIPv6-Diag corpus is the network-flow partition of 1,766,734 records: CIC-IDS-2017 (516,051), CIC-IDS-2018 (546,075) and CIC-DDoS-2019 (704,608), covering five inherited attack families (flooding, brute-force, scan, web-attack and infiltration). The unified benchmark file has 42 columns (27 IPv6 features, a binary label, an is_original flag, flow-key fields and provenance columns). Scope and limits. The 27 IPv6 fields are schema-level features synthesised from RFC-informed distributions, not observations from native IPv6 packet captures. This is a controlled diagnostic substrate for like-for-like comparison of IDS classifiers under synthetic IPv6 semantics and for pipeline and privacy studies; it is not representative of real IPv6 backbone traffic. Against the CAIDA (equinix-chicago) and MAWI (samplepoint-F) traces the median Kolmogorov–Smirnov D is 0.202 and none of 15 evaluated prefix-level features meets D < 0.05. Classifiers trained on it require validation and local adaptation before deployment, and on-corpus rankings are not claimed to transfer. The augmenters are not differentially private; membership-inference risk is disclosed empirically. Auxiliary partition (not part of SynthIPv6-Diag). A host-side partition — CIC-MalMem-2022 (75,546) and EMBER-2024 (158,420), 233,966 rows total — is included only as out-of-domain stress material with statistical IPv6 overlays; it is not network-flow data and is not compared with the network-flow results. The full mixed corpus is 2,000,700 rows. Evaluation protocol. 3× repeated stratified nested cross-validation (outer k = 5, inner k = 3; 15 pooled outer-fold scores), BCa bootstrap 95% confidence intervals (n = 1000), and Holm–Bonferroni-corrected paired Wilcoxon tests. SHA-256 checksums are in CHECKSUMS_SHA256.txt and dataset_metadata.json.



