遇见数据集

Write performance with different stripe counts for Lustre on Joliot Curie (Irene)

收藏
Zenodo2026-04-30 更新2026-05-29 收录
官方服务:

资源简介:

This dataset contains performance measured with the IOR benchmarking tool when writing to the Lustre parallel file system following different access patterns and using different numbers of OSTs (i.e., different Lustre stripe counts). This dataset was used in the experiments reported in [1], in which it was used to train and evaluate the prediction models that select the best stripe count for a given access pattern. A similar dataset, generated on the PlaFRIM platform with the BeeGFS file system, was previously published in [2] and is also used in [1]. All experiments were carried out on the Joliot Curie (Irene) supercomputer at the Très Grand Centre de Calcul (TGCC), operated by CEA (https://www-hpc.cea.fr/), between 2024 and 2025, on the AMD-Rome partition. Each compute node has two 64-core AMD Rome processors and 228 GiB of RAM. The Lustre file system is deployed with 1 MDS+MDT and 40 OSSs, with one OST per OSS, and a default stripe size of 1 MiB. Experiments were generated and orchestrated using the IOPS framework, an open-source tool that automates the description, execution, and post-processing of HPC I/O benchmarking campaigns: https://iops.gitlabpages.inria.fr/ IOR version 4.1.0+dev was used. File format The dataset is provided as a single CSV file in text format. Each row corresponds to one repetition of one experiment configuration. Multiple repetitions of each configuration were executed (at least 3 per configuration), in random order, so the value of the repetition column does not carry any ordering information beyond differentiating between repetitions of the same configuration. Columns nodes: number of compute nodes used. processes_per_node: number of MPI processes per node. procs: total number of MPI processes used (across all nodes), equal to nodes * processes_per_node. filestrategy: either shared-file (a single file is accessed by all processes) or file-per-proc (each process has its own file, IOR option -F). spaciality: either contig (each process accesses a contiguous portion of the file) or strided (1D-strided access pattern, IOR option -s, only with shared-file). reqsize_kb: IOR request size in KiB (IOR option -t). total_data_mb: total amount of data accessed in the experiment, in MiB. The amount accessed per process (for contig patterns) is therefore total_data_mb / procs. ost_number: the Lustre stripe count used, that is, the number of OSTs across which files are striped. On Irene, the stripe count is configured by the user on a per-file basis; in this dataset the requested stripe count was applied to the directory in which IOR creates its files. repetition: differentiates between repetitions of the same configuration. bandwidth_mb: write bandwidth reported by IOR, in MiB/s. time: total write time reported by IOR (including open and close), in seconds. References [1] Francieli Boito, Luan Teylo, Mihail Popov, Laora Aimi, Alexis Bandet, Laércio Lima Pilla, Guillaume Pallez. TOTO: Transparent I/O Tuning for HPC Applications. ACM International Conference on Supercomputing (ICS), July 2026, Belfast, Northern Ireland, United Kingdom. [2] Francieli Boito. Write performance with different numbers of OSTs for BeeGFS in PlaFRIM [Data set]. Zenodo, 2026. https://doi.org/10.5281/zenodo.19890096

提供机构:
Zenodo
创建时间:
2026-04-30
二维码
社区交流群
二维码
科研交流群
商业服务