遇见数据集

Explainable AI for Intrusion Detection on Live Multi-Service Honeypot Telemetry

收藏
Zenodo2026-09-27 更新2026-10-01 收录
官方服务:

资源简介:

# Are Explanations Faithful Under Fire? Evaluating XAI Reliability on Live Honeypot Intrusion-Detection Telemetry ## Supplementary Data and Code — Version 3.0 [![DOI](https://zenodo.org/badge/DOI/10.5281/zenodo.22727709.svg)](https://doi.org/10.5281/zenodo.22727709) > **Authors:** Sridhar G, J B Simha, Rashmi Agarwal> **Affiliation:** School of Computer Science and Application / RACE, REVA University, Bengaluru, India> **Contact:** sridhargovardhan@race.reva.edu.in> **Version:** 3.0 (September 2026) --- ## What Changed from Version 2.0 to 3.0 | Change | v2.0 | v3.0 ||---|---|---|| **Primary faithfulness** | Retraining-based AUDAC | **Frozen-model** masking-based AUDAC (no retraining) || **AUDAC values** | RF 0.679, XGB 0.662, LGB 0.688 | **RF 0.352, XGB 0.348, LGB 0.346, MLP 0.442, LSTM 0.463** || **Flip rate** | Malicious only | **Both malicious AND benign** (trees: 100%/0%) || **Per-instance SHAP-LIME ρ** | Negative (−0.18, buggy) | **Positive** (RF 0.72, XGB 0.43, LGB 0.52, MLP 0.58) || **LOSO** | RF only | **All 4 sklearn models** || **Feature ablation** | dst_port_id only | **4 features** (dst_port_id, src_port_class, hour_of_day, ip_reputation_score) || **Multi-seed** | RF only (3 seeds) | **All 6 models** × 3 seeds || **Service inventory** | 10 services listed | **2,309 unique ports** documented || **Transformer** | Negative control (F1 = 55.9%) | **Convergence failure** (F1 = 0.00 on 2/3 seeds) || **Labels** | "True labels" | **"Operational annotations"** || **Scripts** | 6 scripts | **10 scripts** || **Result files** | 2 CSVs | **7 CSVs + 2 TXT** | --- ## Overview This repository contains the complete reproducibility package for the manuscript *"Are Explanations Faithful Under Fire? Evaluating XAI Reliability on Live Honeypot Intrusion-Detection Telemetry."* The central contribution is a **frozen-model faithfulness evaluation** of model-agnostic explanations (SHAP, LIME) on live adversarial honeypot traffic. Each model is trained once, saved as a frozen checkpoint, and evaluated without retraining — directly measuring whether the trained model relies on the features SHAP identified. --- ## Repository Structure ```.├── README.md│├── # ── Primary Analysis: Frozen-Model Faithfulness ──├── frozen_model_faithfulness.py # Trains, saves, evaluates frozen models (6 models × 3 seeds)├── frozen_fixes.py # Corrects LIME name-matching bug + MLP SHAP Pipeline error├── frozen_faithfulness.csv # Primary results: 18 rows (6 models × 3 seeds)├── frozen_fixes_results.txt # Corrected SHAP-LIME ρ and MLP AUDAC│├── # ── Service Generalisation (Phase 4) ──├── phase4_generalisation.py # LOSO all models, multi-feature ablation, CIs├── loso_all_models.csv # LOSO F1 for RF, XGBoost, LightGBM, MLP (48 rows)├── multi_feature_ablation.csv # 4-feature ablation + combined removal (20 rows)├── per_service_all_models.csv # Per-service F1 for all 4 models (48 rows)├── service_inventory.csv # Complete inventory: 2,309 unique destination ports├── phase4_results.txt # Full Phase 4 output with bootstrap CIs│├── # ── Secondary Analysis: Retraining-Based ──├── run_faithfulness_retraining.py # Retraining-based AUDAC (secondary robustness analysis)├── xai_faithfulness_metrics.py # Faithfulness metrics library (AUDAC, flip rate, sparsity)├── faithfulness_retraining_results.csv # Retraining-based results (6 models, seed=42)│├── # ── Supporting Scripts ──├── per_service_evaluation.py # Per-service F1, LOSO, dst_port_id ablation (RF only, v2)├── bootstrap_ci.py # Bootstrap 95% CIs for AUDAC and flip rate├── reviewer_experiments.py # 6 reviewer-response analyses in one pass├── benchmark_diagnostics.py # Leakage audit and de-duplication analysis├── unsw_corrected_split_rerun.py # Corrected UNSW-NB15 official split re-run│├── # ── Benchmark Results ──└── verified_benchmark_results.csv # UNSW-NB15 and CIC-IDS2017 (6 models)``` --- ## Key Results ### 1. Primary Analysis: Frozen-Model Faithfulness (Table 9) Models trained once → saved as frozen checkpoints (.pkl/.keras) → top-k SHAP features masked with training-set medians → **no retraining**. This directly measures whether the trained model depends on the features SHAP identified. | Model | F1 | Frozen AUDAC ↓ | Mono. | Flip Mal | Flip Ben | Global ρ | Per-inst ρ ||---|---|---|---|---|---|---|---|| RF | 88.6% | **0.352** | Strict | 100% | 0% | 0.87 | 0.72 || XGBoost | 88.2% | **0.348** | Essential | 100% | 0% | 0.93 | 0.43 || LightGBM | 88.1% | **0.346** | Essential | 100% | 0% | 0.85 | 0.52 || MLP | 87.7% | **0.442** | Mostly | 98.8% | 2.2% | 0.87 | 0.58 || LSTM | 83.4% | **0.463** | Mostly | 87.0% | 13.6% | N/A | N/A || Transformer | 67.0% | 0.670 | N/A | 0% | N/A | N/A | N/A | **Monotonicity tiers:** Strict = no violations (k=0–10); Essential = violations only near chance; Mostly = one real violation, overall downward; Anti = accuracy rises (Transformer on retraining — excluded from primary). **Bootstrap 95% CIs (frozen, seed=42):**- AUDAC: RF [0.338, 0.366], XGBoost [0.332, 0.365], LightGBM [0.330, 0.362], MLP [0.425, 0.460]- Flip-rate: RF [1.000, 1.000], XGBoost [1.000, 1.000], LightGBM [1.000, 1.000], MLP [0.970, 1.000] **Key finding:** Tree ensembles flip 100% of malicious predictions but 0% of benign predictions when top SHAP features are masked — confirming that the identified features are specifically diagnostic of malicious behaviour. ### 2. Multi-Seed Variance (seeds 42, 7, 2024) | Model | F1 (mean ± SD) | AUDAC (mean ± SD) ||---|---|---|| RF | 0.891 ± 0.005 | 0.341 ± 0.018 || XGBoost | 0.884 ± 0.003 | 0.428 ± 0.086 || LightGBM | 0.888 ± 0.006 | 0.378 ± 0.027 || MLP | 0.881 ± 0.004 | 0.677 ± 0.024 || LSTM | 0.851 ± 0.014 | 0.461 ± 0.018 || Transformer | 0.223 ± 0.387 | 0.223 ± 0.387 | Transformer achieves F1 = 0.00 on two of three seeds — catastrophic convergence instability on tabular data. ### 3. SHAP-LIME Cross-Method Consistency (corrected in v3) | Model | Global ρ | Per-instance median ρ | IQR ||---|---|---|---|| RF | 0.87 | 0.72 | [0.63, 0.82] || XGBoost | 0.93 | 0.43 | [0.30, 0.59] || LightGBM | 0.85 | 0.52 | [0.38, 0.68] || MLP | 0.87 | 0.58 | [0.48, 0.68] | **Bug fix in v3:** Per-instance values in v2 were negative (median ρ = −0.18) due to a LIME feature-name matching error. LIME with `discretize_continuous=True` returns names like `"payload_length_bytes <= 245.50"` which did not match the original feature names during rank-correlation computation. Corrected in `frozen_fixes.py` using substring matching. ### 4. Service Generalisation (Phase 4 — new in v3) **Service inventory:** 2,309 unique dst_port_id values observed during the 21-day deployment:- 10 deliberately deployed honeypot services (ports 21, 22, 25, 53, 80, 443, 3306, 3389, 5900, 6379) — 651–744 sessions each- 2 additional ports with substantial traffic (8080 HTTP-alt, 8443 HTTPS-alt)- ~2,297 ephemeral scanning/probing ports (1–2 sessions each) **LOSO — all 4 sklearn models (12 services with ≥10 test samples):** | Model | Mean F1 Drop | Services Improved ||---|---|---|| RF | −0.10 pp | 4/12 || XGBoost | +0.54 pp | 4/12 || LightGBM | −0.07 pp | 6/12 || MLP | −0.43 pp | 7/12 | **Multi-feature ablation (all models):** | Feature Removed | RF | XGB | LGB | MLP ||---|---|---|---|---|| dst_port_id | −0.35 pp | −0.51 pp | −0.38 pp | −1.42 pp || src_port_class | +0.61 pp | +0.89 pp | +0.14 pp | +0.63 pp || hour_of_day | +0.81 pp | +1.51 pp | +0.38 pp | +0.96 pp || ip_reputation_score | −0.74 pp | +0.80 pp | +0.25 pp | −0.46 pp || **ALL FOUR** | **+2.80 pp** | **+3.36 pp** | **+2.68 pp** | **+5.43 pp** | Negative drop = F1 *improves* when that feature is removed. Removing dst_port_id improves F1 for all models — no shortcut learning. Removing all four service-related proxies costs only 2.7–5.4 pp, confirming that the core detection signal resides in behavioural features. ### 5. Secondary Analysis: Retraining-Based Faithfulness Retained from v2 as a robustness check. For each removal step k, the model is retrained on remaining features — allowing adaptation, producing higher (more conservative) AUDAC. | Model | F1 | AUDAC ↓ | Flip % | SHAP-LIME ρ ||---|---|---|---|---|| RF | 88.59% | 0.634 | 100% | 0.93 || XGBoost | 88.27% | 0.617 | 98.2% | 0.89 || LightGBM | 88.17% | 0.633 | 100% | 0.86 || MLP | 87.01% | 0.676 | 99.6% | 0.87 || LSTM | 83.15% | 0.630 | 12.8% | 0.93 || Transformer | 59.16% | 0.603 | 15.0% | 0.70 | ### 6. Benchmark Comparison (Table 10) | Model | UNSW-NB15 F1 | UNSW-NB15 AUC | CIC-IDS2017 F1 | CIC-IDS2017 AUC ||---|---|---|---|---|| XGBoost | 92.31% | 98.50% | 99.82% | 99.98% || RF | 92.24% | 98.29% | 99.80% | 99.98% || LightGBM | 92.06% | 98.49% | 99.90% | 99.99% || LSTM | 91.32% | 98.03% | 98.45% | 99.96% || MLP | 91.04% | 97.96% | 98.84% | 99.94% || Transformer | 90.77% | 96.86% | 97.27% | 99.92% | Note: different feature schemas (42 features for UNSW-NB15, 78 for CIC-IDS2017 vs 13 for the honeypot), different label definitions, and different preprocessing pipelines prevent direct comparison. Lower live-traffic scores may reflect adversarial difficulty but should not be interpreted as a causal relationship. --- ## Script Descriptions ### `frozen_model_faithfulness.py` — Primary Faithfulness (NEW in v3) The central experiment addressing the peer-review requirement for frozen-model evaluation. For each of 6 models × 3 seeds:1. Trains the model on the 85% training partition (n = 9,265)2. Saves a frozen checkpoint to `frozen_models/` (.pkl for sklearn, .keras for TensorFlow)3. Loads the frozen checkpoint for all subsequent evaluations4. Computes SHAP values (TreeExplainer for trees, KernelExplainer for deep models)5. Runs masking-based AUDAC (k = 0, 1, ..., 13) — **no retraining**6. Computes flip rate for **both** malicious and benign sessions (n = 500 each)7. Reports bootstrap 95% CIs (500 iterations) **Data flow:**```Raw data (10,900) → 15% held-out (n=1,635, seed=42) → Train on 9,265 → Save frozen checkpoint → SHAP (frozen model, bg=100 k-means) → Feature masking (frozen model) → AUDAC + flip rate``` **Runtime:** ~16 hours (3 seeds × 6 models). Tree models ~5 min each; deep models ~2 hours each (KernelExplainer). ### `frozen_fixes.py` — LIME Bug Fix + MLP SHAP Fix (NEW in v3) Loads seed=42 frozen models and corrects two issues:1. **LIME feature-name matching:** LIME with `discretize_continuous=True` returns discretised names (e.g., `"payload_length_bytes <= 245.50"`) that did not match original feature names. Fixed with substring matching.2. **MLP SHAP:** sklearn Pipeline's `feature_names_in_` property rejected KernelExplainer input. Fixed with a DataFrame wrapper that bypasses validation. Also computes per-instance SHAP-LIME rank correlation (median + IQR) for RF, XGBoost, LightGBM, MLP. **Runtime:** ~30 minutes. ### `phase4_generalisation.py` — Service Generalisation (NEW in v3) Addresses the reviewer requirement for LOSO on all models and broader feature ablation:1. **Service inventory:** maps all 2,309 destination ports with sample sizes and class distributions2. **Per-service F1:** held-out F1 for each of 12 services (≥10 test samples) × 4 models3. **LOSO:** leave-one-service-out for RF, XGBoost, LightGBM, MLP4. **Multi-feature ablation:** removes dst_port_id, src_port_class, hour_of_day, ip_reputation_score individually and together5. **Bootstrap CIs** for per-service and LOSO F1 **Runtime:** ~25 minutes. ### `run_faithfulness_retraining.py` — Secondary Robustness Analysis The v2 retraining-based AUDAC analysis, now labelled as secondary. For each feature-removal step, the model is retrained on remaining features. ### `xai_faithfulness_metrics.py` — Faithfulness Metrics Library Implements AUDAC (14-point dense curve), sparsity, stability, and perturbation flip rate following Arreche et al. [8] and Gaspar et al. [10]. Also computes MCC, balanced accuracy, and SHAP-LIME Spearman rank correlation. ### `reviewer_experiments.py` — Combined Reviewer-Response Experiments Six analyses: LIME sufficiency (2K vs 5K), Mahalanobis distance of perturbed samples (99.5% within 95th-percentile envelope), Shapiro-Wilk normality test (W = 0.958, p = 0.76), per-instance SHAP-LIME ρ, multi-seed RF variance, bootstrap CIs. ### `benchmark_diagnostics.py` — Leakage Audit Single-feature ROC-AUC (max 0.744 CIC, 0.783 UNSW — no identifier leakage), duplicate analysis (CIC 10.9% duplicates, < 0.4 pp F1 impact). ### `unsw_corrected_split_rerun.py` — Corrected UNSW Split Re-runs XGBoost, LSTM, Transformer on the official UNSW-NB15 partition with a built-in swap guard. --- ## Datasets | Dataset | Source | Records | Usage ||---|---|---|---|| Honeypot corpus | This work | 10,900 sessions (13 features + label) | Primary evaluation (Tables 4–9) || UNSW-NB15 | [UNSW Sydney](https://research.unsw.edu.au/projects/unsw-nb15-dataset) | 82,332 / 175,341 | Benchmark (Table 10) || CIC-IDS2017 | [UNB CIC](https://www.unb.ca/cic/datasets/ids-2017.html) | 2,830,743 flows | Benchmark (Table 10) | The honeypot dataset (`honeypot_features_10900.csv`) contains 15 columns: `row_id` (drop before training), 13 numeric features, and `label` (binary 0/1). Train/test: `train_test_split(X, y, test_size=0.15, stratify=y, random_state=42)`. ### Three-Stage Labelling Protocol Labels are termed **operational annotations** to acknowledge that the labelling process uses some of the same observables that constitute the feature set. - **Stage 1:** Automated pre-labelling using rule-based signatures on raw log fields — **no access to engineered features**- **Stage 2:** Entropy-based anomaly scoring using payload_entropy and auth_failure_ratio- **Stage 3:** Expert review — both annotators (5+ years SOC experience) independently examined session metadata, log content, and the full 13-feature summary table before consensus discussion Inter-rater reliability: Cohen's κ = 0.91. Disagreement rate: 8.3%, resolved by consensus. --- ## Evaluation Protocol ### Primary: Frozen-Model Faithfulness (Table 9)```Raw data (10,900) → 15% held-out (n=1,635, seed=42) → Train on 9,265 → Save frozen checkpoint → Held-out evaluation (Table 7) → SHAP computation (frozen) → Feature masking on frozen model (no retraining) → Table 9: AUDAC, flip rate (malicious + benign), SHAP-LIME ρ```- SHAP: TreeExplainer (RF, XGBoost, LightGBM); KernelExplainer (100 k-means background) for MLP, LSTM, Transformer- LIME: n_samples = 2,000, discretize_continuous = True, seed = 42- Flip rate: n = 500 correctly classified sessions per class, top-10 features masked- Bootstrap: 500 iterations for AUDAC and flip-rate 95% CIs ### Secondary: Retraining-Based (Supplementary)Same split; model retrained at each removal step. Produces higher (more conservative) AUDAC. ### Statistical Tests- Paired Wilcoxon signed-rank (10 fold-level F1, dependent observations — fold dependence acknowledged as limitation; p-values approximate)- Holm-Bonferroni for 5 pairwise comparisons vs RF- Cohen's D effect sizes: XGB 0.08, LGB 0.48, TRF 0.45, LSTM 1.32, MLP 1.56- Shapiro-Wilk: W = 0.958, p = 0.76 ### Benchmarks (Table 10)- UNSW-NB15: official split (82,332 / 175,341); 42 features; binary target- CIC-IDS2017: 70/30 stratified split, three seeds; 78 features --- ## Reproducibility ```Python >= 3.10scikit-learn 1.3.2XGBoost 2.0.3LightGBM 4.1.0TensorFlow 2.21.0SHAP 0.44.1LIME 0.2.0.1Random seed 42 (primary); 7, 2024 (multi-seed)``` ```bashpip install pandas numpy scikit-learn scipy xgboost lightgbm shap lime tensorflow``` ### Quick Start ```bash# 1. Primary: frozen-model faithfulness (6 models × 3 seeds, ~16 hrs)python frozen_model_faithfulness.py # 2. Fix LIME + MLP SHAP (loads frozen models, ~30 min)python frozen_fixes.py # 3. Service generalisation: LOSO all models + ablation (~25 min)python phase4_generalisation.py # 4. Secondary: retraining-based AUDAC (~2 hrs)python run_faithfulness_retraining.py # 5. Reviewer-response experiments (~45 min)python reviewer_experiments.py # 6. Benchmark diagnosticspython benchmark_diagnostics.py``` --- ## Version History | Version | Date | Changes ||---|---|---|| 1.0 | June 2025 | Initial deposit: retraining-based faithfulness, 3-point AUDAC || 2.0 | February 2026 | Dense 14-point AUDAC, per-service evaluation, bootstrap CIs, reviewer experiments, benchmark diagnostics || **3.0** | **September 2026** | **Frozen-model faithfulness (primary). LIME bug fix (per-instance ρ corrected). MLP SHAP fix. Phase 4: LOSO all models, 4-feature ablation, 2,309-port inventory. Multi-seed all 6 models. Benign flip rates. "Operational annotations" terminology. Retraining moved to secondary.** | --- ## Citation ```bibtex@article{sridhar2026xai_honeypot, title = {Are Explanations Faithful Under Fire? {E}valuating {XAI} Reliability on Live Honeypot Intrusion-Detection Telemetry}, author = {Sridhar, G. and Simha, J. B. and Agarwal, Rashmi}, year = {2026}, journal = {Frontiers in Computer Science -- Computer Security}, doi = {10.5281/zenodo.22727709}}``` --- ## License Code: [MIT License](https://opensource.org/licenses/MIT).Dataset: [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/).

提供机构:
Zenodo
创建时间:
2026-09-27
二维码
社区交流群
二维码
科研交流群
商业服务