RevealTelemetryDatasetforMLInfraProfilingAnomalyDetection
收藏资源简介:
Reveal是一个大规模的、经过精心策划的硬件遥测数据集,收集自运行多样化机器学习工作负载的高性能计算环境。该数据集使得可以对系统级分析、无监督异常检测和机器学习基础架构优化进行可重现的研究。它包括低级硬件和操作系统指标,这些指标对操作员完全可见,允许进行异常检测而不需要工作负载知识或仪器化。
Reveal is a large-scale, curated hardware telemetry dataset collected from high-performance computing environments running diverse machine learning workloads. This dataset enables reproducible research for system-level analysis, unsupervised anomaly detection, and machine learning infrastructure optimization. It includes low-level hardware and operating system metrics that are fully visible to operators, allowing anomaly detection without requiring workload knowledge or instrumentation.
Reveal: Hardware Telemetry Dataset for Machine Learning Infrastructure Profiling and Anomaly Detection
数据集详情
数据集描述
- Reveal是一个大规模、经过整理的高性能计算硬件遥测数据集,收集自运行多样化机器学习工作负载的环境
- 支持系统级分析、无监督异常检测和机器学习基础设施优化的可重复研究
- 数据集配套论文《Detecting Anomalies in Systems for AI Using Hardware Telemetry》(Chen等,牛津大学,2025年)
基本信息
- 策划者:Ziji Chen, Steven W. D. Chien, Peng Qian, Noa Zilberman(牛津大学工程科学系)
- 共享者:Ziji Chen(联系方式:ziji.chen@eng.ox.ac.uk)
- 语言:英语(元数据和文档)
- 许可证:CC BY 4.0
数据集来源
- 论文:https://arxiv.org/abs/submit/6934461
- DOI:https://doi.org/10.5281/zenodo.17470313
用途
直接用途
- 系统遥测中的无监督异常检测研究
- 硬件指标的多元时间序列建模
- 跨子系统交互研究(CPU、GPU、内存、网络、存储)
- 开发性能感知的机器学习基础设施工具
- AIOps和ML系统健康监控的异常检测模型训练或基准测试
超出范围用途
- 推断或重建用户工作负载或模型行为
- 最终用户应用程序性能基准测试
- 涉及个人、机密或专有数据重建的任何用途
数据集结构
核心字段
timestamp:样本的UTC时间host_id:主机或节点标识符metric_name:测量计数器的名称value:记录的数值subsystem:子系统类别(CPU、GPU、Memory、Network、Storage)
数据集创建
数据收集与处理
- 收集工具:perf、procfs、nvidia-smi和标准Linux实用程序
- 采样间隔:100毫秒
- 数据规模:每个主机约150种原始指标类型,扩展为约700个时间序列通道
工作负载和系统
- 工作负载:超过30种机器学习应用程序(BERT、BART、ResNet、ViT、VGG、DeepSeek、LLaMA、Mistral)
- 数据集:GLUE/SST2、WikiSQL、PASCAL VOC、CIFAR、MNIST
- 系统配置:双节点GPU HPC集群(NVIDIA V100和H100、Intel Xeon CPU、InfiniBand HDR100)
数据生产者
所有数据由作者在受控环境中使用合成工作负载生成,不包含用户或私人信息。
引用
bibtex @misc{chen2025detectinganomaliesmachinelearning, title={Detecting Anomalies in Machine Learning Infrastructure via Hardware Telemetry}, author={Ziji Chen and Steven W. D. Chien and Peng Qian and Noa Zilberman}, year={2025}, eprint={2510.26008}, archivePrefix={arXiv}, primaryClass={cs.PF}, url={https://arxiv.org/abs/2510.26008}, }




