Reveal: Hardware Telemetry Dataset for Machine Learning Infrastructure Profiling and Anomaly Detection
收藏资源简介:
This dataset accompanies the paper “Detecting Anomalies in Machine Learning Infrastructure via Hardware Telemetry” (Chen et al., University of Oxford, 2025).It contains host-level time-series telemetry from a diverse set of machine learning (ML) applications, each executed ten times under controlled conditions on GPU-based HPC systems. Each run includes over 150 types of system-level metrics—covering CPU, GPU, memory, network, and storage subsystems—sampled at 100 ms intervals.The dataset also provides metadata including workload type (LLM or non-LLM), workload phase (training, fine-tuning, or inference), task, model architecture, and host information. This dataset supports reproducible research in ML systems, performance profiling, and unsupervised anomaly detection, and forms the basis of the Reveal framework described in the associated paper.



