遇见数据集

OpenTelemetry AIOps Benchmark Dataset: ML-based Anomaly Detection across Traces, Metrics, and Logs

收藏
Zenodo2026-04-07 更新2026-05-26 收录
官方服务:

资源简介:

Companion dataset for the paper "Evaluating ML-based anomaly detection on unified OpenTelemetry telemetry: an empirical study across traces, metrics, and logs" submitted to IEEE Access. This dataset contains processed telemetry features and complete experiment results from a chaos engineering study on a microservices testbed (Weaveworks Sock Shop, 13 services) instrumented with OpenTelemetry and deployed on AWS EKS (Kubernetes 1.29). Experiment Design A 24-hour baseline of normal operation was collected, followed by a fault injection campaign of 40 injections (8 fault types x 5 repetitions) using Chaos Mesh. Fault types: network delay, pod kill, CPU stress, memory stress, disk fill, HTTP abort, container kill, and pod failure. Each injection lasted 5 minutes with 10-minute recovery windows. Telemetry was collected across three OpenTelemetry signals: distributed traces (Jaeger), infrastructure metrics (Prometheus), and application logs (Loki). ML Models Evaluated Six anomaly detection models were benchmarked: Isolation Forest, One-Class SVM, LSTM Autoencoder, Transformer Autoencoder, 1D-CNN Autoencoder, and a supervised Random Forest baseline. All models were evaluated using 5-fold cross-validation with 1,800 Optuna hyperparameter optimization trials. A cooldown exclusion methodology was applied to avoid penalizing models during post-fault recovery periods. Key Findings Transformer Autoencoder on traces achieved the best unsupervised AUC-ROC (0.800) Isolation Forest max-fusion across signals achieved F1 of 0.752 Feature ablation showed mean-only metrics yielded AUC 0.964, outperforming full feature sets Supervised Random Forest on metrics achieved F1 0.848 and AUC 0.985 Dataset Contents features/: Processed Parquet feature files (metrics, logs, traces, combined) with 60-second aggregation windows results/: All experiment results including per-model/signal/fold metrics, raw anomaly scores (.npy), fusion results, statistical significance tests, prevalence sensitivity analysis, episode-level analysis, feature ablation, and supervised baseline chaos-experiments/: Chaos Mesh fault injection YAML configurations scripts/: Fault injection campaign orchestration code ml-pipeline/: Preprocessing and training pipeline code analysis/: Post-hoc analysis scripts (significance tests, prevalence sensitivity, feature ablation) figures/: Figure generation scripts for all paper visualizations

提供机构:
Zenodo
创建时间:
2026-04-07
二维码
社区交流群
二维码
科研交流群
商业服务