遇见数据集

Streaming Pipeline Anomaly Detection Benchmark: Kafka and Flink Telemetry Dataset

收藏
Zenodo2026-05-01 更新2026-05-26 收录
官方服务:

资源简介:

This deposit is the open benchmark accompanying "ML-based anomaly detection for streaming pipelines: an empirical benchmark with cross-workload generalization on Apache Kafka and Flink" (Anjum, 2026; submitted to Elsevier Array). This is v3 of the deposit. It supersedes earlier versions (5,000 vs 500 events/s pilot, six-model panel) which described the initial single-workload study; the current manuscript describes the two-workload eight-model study captured here. What is in this deposit Two independent fault-injection campaigns on Apache Kafka 3.9 (KRaft) and Apache Flink 1.19, each deployed on a six-node Amazon EKS 1.30 cluster: E-commerce workload at a nominal 5,000 events/s (±20% sinusoidal variation), 100k product catalogue, three-stage stateful Flink topology (event enrichment → windowed aggregation → fraud detection). Industrial IoT workload at 10,000 events/s from 100,000 simulated devices across 500 geographic locations, three-stage Flink topology (location enrichment → per-location aggregation → temperature/battery alerting). Both workloads use the same fault taxonomy (eight streaming-specific faults F1–F8: broker crash, network partition, disk I/O saturation, consumer lag spike, rebalance storm, poison pill, checkpoint failure, OOM exhaustion), the same Chaos Mesh injection protocol, and the same telemetry/feature pipeline. Each workload contributes 80 controlled fault-injection runs (8 faults × 10 repetitions) plus a 24-hour normal baseline; total 160 runs. Files Path Size Description ecommerce/baseline_24h.parquet ~17 MB 24-hour pre-campaign baseline ecommerce/campaign.parquet ~56 MB 80-run fault campaign ecommerce/campaign_manifest.json ~180 KB Per-run timestamps, labels, recovery status ecommerce/features/features_{15,30,60,120}s.parquet per-window Engineered features after variance + correlation filtering iot/baseline_24h_iot.parquet ~17 MB IoT 24h baseline iot/campaign_iot.parquet ~56 MB IoT 80-run campaign iot/campaign_manifest_iot.json ~180 KB IoT manifest iot/features/features_{15,30,60,120,300}s.parquet per-window IoT features iot/baseline_stats_iot.json ~590 KB Per-feature mean/std/min/max/p50/p95 analysis/cross_workload/*.csv small Phase-7.7 cross-workload outputs (panel_overall, panel_per_fault, subexp1–4 results, mixed-effects, curated-rule baseline per-fold) code/sagemaker/preprocessing.py, training.py, post_hoc_analysis.py small ML pipeline scripts code/infrastructure/ small Terraform IaC for EKS + Helm values code/streaming-app/ small Avro schemas + Flink jobs (Java) + Python load generators code/analysis/ small Cross-workload analysis scripts (mixed-effects, curated rules) Models evaluated Eight semi-supervised: Isolation Forest, One-Class SVM, LSTM-AE, Transformer-AE, 1D-CNN-AE, LSTM-VAE, Deep SVDD, DAGMM. Plus a supervised Random Forest reference, a curated production-style rule baseline (Burrow/Xinfra/Ververica-style: consumer lag, failed checkpoints, backpressure, GC rate), and a statistical max-z diagnostic. The full eight-model panel is evaluated at the 30-second anchor on both workloads; IoT additionally covers the 15s/60s/120s/300s sensitivity grid with the full panel, while e-commerce covers 15s/60s/120s with the original six-model panel. The asymmetry is documented in the manuscript's Threats to Validity section. Reproducibility All infrastructure is provisioned via Terraform; Kafka, Flink, and Chaos Mesh deployments are codified in Helm and Kubernetes manifests. The fault-injection campaign is automated end-to-end by a Python orchestrator. Training is containerised on AWS EC2 g5.2xlarge (NVIDIA A10G GPU) for deep models and c5.4xlarge for classical models. Random seeds are fixed at 42 + fold_id. The 10-fold cross-validation uses a four-role split (train/hpo_val/cal/test) to avoid validation triple-use. The companion code repository is at https://github.com/mateenali66/streaming-anomaly-detection. Citation Anjum, M. A. (2026). ML-based anomaly detection for streaming pipelines: e-commerce and industrial IoT campaigns on Apache Kafka and Flink (v3 dataset) Changelog v3 (current): Adds IoT workload (10,000 eps, 100k devices), Deep SVDD and DAGMM models, 300-second window, mixed-effects RQ2 analysis, curated production-style rule baseline, full cross-workload analysis (bucket consistency, per-fault Spearman, optimal window transferability, cross-train transfer). v2: Added e-commerce phase B campaign (faults F5–F8) to the original phase A (F1–F4). v1: Initial e-commerce campaign at 5,000 eps with six semi-supervised models.

提供机构:
Zenodo
创建时间:
2026-05-01
二维码
社区交流群
二维码
科研交流群
商业服务