遇见数据集

HPC-ODA Dataset Collection

收藏
Zenodo2021-04-08 更新2026-05-25 收录
数据链接:
官方服务:

资源简介:

HPC-ODA is a collection of datasets acquired on production HPC systems, which are representative of several real-world use cases in the field of Operational Data Analytics (ODA) for the improvement of reliability and energy efficiency. The datasets are composed of monitoring sensor data, acquired from the components of different HPC systems depending on the specific use case. Two tools, whose overhead is proven to be very light, were used to acquire data in HPC-ODA: these are the DCDB and LDMS monitoring frameworks. The aim of HPC-ODA is to provide several vertical slices (here named segments) of the monitoring data available in a large-scale HPC installation. The segments all have different granularities, in terms of data sources and time scale, and provide several use cases on which models and approaches to data processing can be evaluated. While having a production dataset from a whole HPC system - from the infrastructure down to the CPU core level - at a fine time granularity would be ideal, this is often not feasible due to the confidentiality of the data, as well as the sheer amount of storage space required. HPC-ODA includes 6 different segments: Power Consumption Prediction: a fine-granularity dataset that was collected from a single compute node in a HPC system. It contains both node-level data as well as per-CPU core metrics, and can be used to perform regression tasks such as power consumption prediction. Fault Detection: a medium-granularity dataset that was collected from a single compute node while it was subjected to fault injection. It contains only node-level data, as well as the labels for both the applications and faults being executed on the HPC node in time. This dataset can be used to perform fault classification. Application Classification: a medium-granularity dataset that was collected from 16 compute nodes in a HPC system while running different parallel MPI applications. Data is at the compute node level, separated for each of them, and is paired with the labels of the applications being executed. This dataset can be used for tasks such as application classification. Infrastructure Management: a coarse-granularity dataset containing cluster-wide data from a HPC system, about its warm water cooling system as well as power consumption. The data is at the rack level, and can be used for regression tasks such as outlet water temperature or removed heat prediction. Cross-architecture: a medium-granularity dataset that is a variant of the Application Classification one, and shares the same ODA use case. Here, however, single-node configurations of the applications were executed on three different compute node types with different CPU architectures. This dataset can be used to perform cross-architecture application classification, or performance comparison studies. DEEP-EST Dataset: this medium-granularity dataset was collected on the modular DEEP-EST HPC system and consists of three parts.These were collected on 16 compute nodes each, while running several MPI applications under different warm-water cooling configurations. This dataset can be used for CPU and GPU temperature prediction, or for thermal characterization. The HPC-ODA dataset collection includes a readme document containing all necessary usage information, as well as a lightweight Python framework to carry out the ODA tasks described for each dataset.

HPC-ODA是一个采集自生产级高性能计算(High Performance Computing, HPC)系统的数据集集合,其涵盖了运维数据分析(Operational Data Analytics, ODA)领域中多个用于提升系统可靠性与能源效率的真实应用场景。该数据集由监控传感器数据组成,这些数据依据具体应用场景,从不同HPC系统的组件中采集得到。本次数据集采集过程中使用了两款经证实开销极低的监控工具:DCDB与LDMS监控框架。HPC-ODA的目标是为大规模HPC部署中的监控数据提供多个垂直切片(本文中称为分段)。这些分段在数据源与时间尺度上具备不同的粒度,且提供了多个可用于评估数据处理模型与方法的应用场景。尽管从基础设施到CPU核心级别的全层级、细时间粒度的HPC系统生产数据集堪称理想,但由于数据保密性与海量存储需求,这一目标通常难以实现。HPC-ODA包含6个不同的分段: 1. 功耗预测(Power Consumption Prediction):细粒度数据集,采集自某HPC系统中的单个计算节点,包含节点级数据与每CPU核心指标,可用于执行功耗预测等回归任务。 2. 故障检测(Fault Detection):中粒度数据集,采集自某遭受故障注入测试的单个计算节点,仅包含节点级数据,以及对应节点上实时运行的应用与故障的标签,可用于执行故障分类任务。 3. 应用分类(Application Classification):中粒度数据集,采集自某HPC系统中的16个计算节点,这些节点运行着不同的并行消息传递接口(Message Passing Interface, MPI)应用。数据以计算节点为单位进行分隔,且与正在运行的应用标签相关联,可用于应用分类等任务。 4. 基础设施管理(Infrastructure Management):粗粒度数据集,包含某HPC系统的集群级数据,涉及温水冷却系统与功耗情况。数据以机架为单位,可用于执行出水口水温、余热预测等回归任务。 5. 跨架构(Cross-architecture):中粒度数据集,为应用分类数据集的变体,共享相同的ODA应用场景。但本数据集的应用以单节点配置的形式,在三种具备不同CPU架构的计算节点上运行,可用于执行跨架构应用分类或性能对比研究。 6. DEEP-EST数据集:该中粒度数据集采集自模块化DEEP-EST HPC系统,包含三个子部分。每个子部分均采集自16个计算节点,在不同温水冷却配置下运行多款MPI应用。该数据集可用于CPU与GPU温度预测,或热特性表征。 HPC-ODA数据集集合附带一份包含所有必要使用说明的README文档,以及一个轻量级Python框架,用于完成各数据集对应的ODA任务。

提供机构:
Zenodo
创建时间:
2020-09-16
二维码
社区交流群
二维码
科研交流群
商业服务