遇见数据集

SynthIPv6-Diag: A Controlled, RFC-Informed Synthetic IPv6 Network-Flow Corpus for IDS Benchmarking and Membership-Inference Auditing

收藏
Zenodo2026-08-18 更新2026-08-20 收录
官方服务:

资源简介:

SynthIPv6-Diag is a controlled, RFC-informed synthetic IPv6 network-flow corpus for reproducible intrusion-detection benchmarking and membership-inference auditing. It is derived from three public IPv4 IDS corpora—CIC-IDS-2017, CIC-IDS-2018 and CIC-DDoS-2019—whose flow records are deterministically mapped to configured IPv6 addressing and transition templates (6to4, NAT64, Teredo and modified EUI-64). The resulting records contain 27 engineered IPv6 header and address-dynamics fields governed by implementation-level relational and range constraints. The headline network-flow corpus contains 1,766,734 records: - CIC-IDS-2017: 516,051 records- CIC-IDS-2018: 546,075 records- CIC-DDoS-2019: 704,608 records It covers five inherited attack families: flooding, brute-force, scan, web attack and infiltration. The unified network view contains 42 columns: 27 IPv6-derived fields, a binary label, an originality flag, flow-key fields and provenance fields. Version 2.2 adds a machine-readable protocol-constraint manifest with 21 verified implementation-level correction labels, a data-use limitations note, a row-indexed provenance/usage index, and version-specific SHA-256 checksums. Correction counters describe rule actions rather than unique records; one row may trigger more than one correction label. Important data-use limitation: 283,317 legacy generated attack rows (16.04% of the network corpus and 28.61% of attack-labelled rows) have `binary_label=1` and `is_original=0`, but their original subtype and generator-method provenance were not retained in version 2.1. No missing subtype is imputed in version 2.2. `PROVENANCE_USAGE_INDEX_v2.2.csv.gz` identifies these rows by `network_row_index` and marks them as `multiclass_eligible=false`. They are designated for binary benign-versus-attack experiments only and must be excluded from multi-class, attack-family and per-subtype analyses. For backward compatibility, the original `generation_method` field is preserved. Its legacy value `zero_day_evolution` is clarified in the v2.2 usage index as `synthetic_challenge_variant`: it denotes a synthetic perturbation of known attack data, not an observed real-world zero-day exploit or a held-out empirical attack family. The archive also includes an auxiliary host-side partition derived from CIC-MalMem-2022 and EMBER2024. This 233,966-row partition is explicitly out-of-domain stress material; it is not network-flow data and is not included in the headline network-flow counts, network-fidelity summaries or network-classifier comparisons. The full mixed archive contains 2,000,700 rows. Scope and limitations: the 27 IPv6 fields are schema-level synthetic features, not observations from native IPv6 packet captures. External validation compares 15 shared IPv6 header fields against authorised CAIDA and MAWI trace-derived packet-header tables. It is a cross-level marginal diagnostic, not flow-level validation: packet-to-flow aggregation would require trace-specific timeout and direction choices not shared with this corpus. For the selected comparison, none of 15 fields meets the pre-specified KS D < 0.05 criterion and the median KS D is 0.202. These values are not lower or upper bounds on unavailable flow-level realism. SynthIPv6-Diag is therefore suitable for controlled benchmarking, pipeline ablation and privacy auditing, not as a representative IPv6 backbone sample or a deployment-ready training set. The evaluation protocol uses 3 repeated stratified nested cross-validation runs with 5 outer folds and 3 inner folds, 1,000-resample bootstrap confidence intervals, and Holm–Bonferroni-corrected paired Wilcoxon tests. Random Forest, XGBoost, LightGBM and logistic regression are used for nested comparison; MLP and CNN-LSTM are supplementary detection diagnostics. The fixed Random Forest Train-Synthetic Test-Real procedure is a separate synthetic-to-real utility diagnostic. The public release contains the corpus, metadata and provenance manifests, documentation, the v2.2 protocol-constraint manifest, the provenance/usage index and checksums. CAIDA traces and CAIDA-derived feature tables are not redistributed. The dataset is released under CC BY 4.0. Version DOI: 10.5281/zenodo.21999352 Concept DOI: 10.5281/zenodo.19503445

提供机构:
Zenodo
创建时间:
2026-08-18
二维码
社区交流群
二维码
科研交流群
商业服务