GWDG GPU Node Telemetry Dataset for Observability-Aware Early Warning of GPU Detachment Failures (2025-2026)
收藏资源简介:
This dataset provides a large-scale, structured collection of sanitized time-series telemetry from production GPU-equipped HPC nodes at the GWDG high-performance computing infrastructure. It is aligned with operator-curated GPU failure incidents observed between January 2025 and February 2026, supporting reproducible research on observability-aware early warning of GPU instability. The dataset, titled “GWDG GPU Node Telemetry Dataset for Observability-Aware Early Warning of GPU Detachment Failures (2025–2026)” (DOI: 10.5281/zenodo.19052367, Version 1.0.0), contains: 21 incident and reference telemetry windows 21 aligned telemetry files and 21 metadata files 7 unique GPU nodes 15,189,625 total telemetry rows Time coverage from 2025-01-01 00:00:00 UTC to 2026-02-09 15:30:00 UTC Each telemetry window contains on average 723,315 rows and follows a fully consistent schema (11 columns) across all files, ensuring reproducibility and compatibility with machine learning workflows. Data Structure and Observability Dimensions Each telemetry record captures multiple operational dimensions: Temporal dimension: timeUtc for time-series alignment Node and hardware context: node, gpu, uuid, device Metric definition: metric, value Workload context: job (linking telemetry to scheduler activity) Monitoring infrastructure: instance (exporter/scrape origin) Hardware/software configuration: modelName, driverVersion This structure enables multi-dimensional analysis of GPU behavior across hardware, workload, and monitoring layers. Observability-Aware Failure Modeling Unlike traditional datasets that focus solely on numeric anomalies, this dataset explicitly supports failure modes where observability degrades. These include: metric disappearance scrape instability and failures exporter or pipeline inconsistencies structural monitoring gaps Such conditions are particularly relevant for GPU detachment and bus-related failures, where GPUs become inaccessible without clear numeric precursors. Incident Coverage The dataset includes 16 labeled failure windows spanning recurring GPU-related issues, including: GPU detachment and PCIe/bus-related failures general GPU malfunctions GPU loss or disappearance uncertain or potentially non-GPU-related anomalies In addition, the dataset provides quarterly reference windows representing healthy system behavior: gpus_when_good_2025-Q1 gpus_when_good_2025-Q2 gpus_when_good_2025-Q3 gpus_when_good_2025-Q4 gpus_when_good_2026-Q1 These reference periods capture expected GPU availability under normal conditions, including scheduler-aware states (e.g., maintenance, draining, health-check failures), enabling robust benchmarking of anomaly detection and early warning methods. Observability Sources The telemetry is derived from Prometheus monitoring archives and integrates multiple observability planes: GPU telemetry via NVIDIA DCGM Node and OS telemetry via Prometheus node exporter Monitoring pipeline telemetry (e.g., scrape success, scrape duration) Scheduler state signals derived from Slurm exporter metrics Each incident window includes aligned pre-incident and post-incident intervals, preserving temporal context around failures. Data Sanitization and Privacy All identifiers, including node hostnames, GPU UUIDs, and exporter instances, were pseudonymized using deterministic salted hashing. The dataset was further sanitized to remove: filesystem paths environment references URLs and IP addresses sensitive textual patterns The release contains no user data, credentials, network topology information, or application-level traces. Research Applications With more than 15.1 million telemetry rows, consistent schema, and aligned incident metadata, this dataset provides a foundation for: observability-aware anomaly detection early warning systems for GPU detachment failures predictive maintenance in GPU-enabled HPC systems benchmarking reliability and monitoring frameworks Related Work This dataset accompanies the study: Bidollahkhani, M., Nordsiek, F., & Kunkel, J. M. (2026).“When GPUs Fail Quietly: Observability-Aware Early Warning Beyond Numeric Telemetry.”arXiv: https://arxiv.org/abs/2603.28781



