Streaming Pipeline Anomaly Detection Benchmark: Kafka and Flink Telemetry Dataset
收藏资源简介:
This deposit accompanies the manuscript submitted to Elsevier Array: "ML-based anomaly detection for streaming pipelines: an empirical benchmark with cross-workload generalization on Apache Kafka and Flink". It contains two complete fault-injection campaigns (e-commerce and industrial IoT) on Apache Kafka 3.9 (KRaft) and Apache Flink 1.19, plus engineered features, analysis outputs, and the code that produced them. This is v8. The campaign datasets are unchanged from v6. v8 revises the analysis layer and adds: Temporal-contiguity audit and a corrected sequence builder, with a three-arm ablation separating training-time from evaluation-time effects of sequences that span cooldown gaps and cross-run boundaries. A matched point-versus-sequence comparison in which both model families are scored on an identical window set at identical prevalence. A chronological operational replay of the alert stream, reporting alerting duty cycle and episode structure rather than analytically derived false-alert rates. A full-panel fixed-FPR latency analysis covering all ten models, and a score-saturation diagnostic explaining why several models cannot meet any low false-positive budget. A corrected cross-workload transfer experiment. The previous analysis standardised the target workload with source-workload statistics, which saturated the score function and produced constant score vectors; the corrected analysis refits standardisation on the target and is fold-structured with clustered statistics. Promotion of the EWMA classical baseline into the main model comparison. A symptom-onset analysis checking that injection-period labels coincide with measurable system effects. A per-fold feature-selection sensitivity analysis. verify_claims.py, a harness that recomputes every numeric claim in the manuscript from these artifacts and reports pass or fail. Deposit contents are data and code only. Licence CC BY 4.0.



