遇见数据集

Profiling results for electron-positron plus jets production with Pepper

收藏
Zenodo2024-08-09 更新2026-05-26 收录
官方服务:

资源简介:

Profiling results for electron positron plus jets production with Pepper v1.1.1 Contents This includes data/output for the following GPU accelerated unweighted event generation runs: A run over a few seconds to produce pepper-internal timing results (`*timers.csv` files) The same run, but with `nvprof` (`*.nvprof` files) A run with a single event batch, with `ncu --import-source on --set full` (`*.ncu-rep` files) The simulated process is electron-positron pair production with n jets, with n = 0...4, at 13 TeV. The configuration files (`pepper.ini`) are included. Also the `pepper_cache` directories are included. All results are for the "main" pepper variant (which uses Kokkos to utilise the GPU) and for the "native" pepper variant (which uses CUDA directly). Software stack and CPU/GPU hardware details The software/hardware used for the profiling are as follows: GPU Driver: NVIDIA-SMI: 550.100, Driver Version: 550.100, CUDA Version: 12.4 GPU: Tesla V100S-PCIE-32GB Cuda compilation tools: release 11.6, V11.6.124 Kokkos 4.3.01 gcc (GCC) 11.4.1 20231218 (Red Hat 11.4.1-3) Intel(R) Xeon(R) Silver 4214R CPU @ 2.40GHz Note that the event generation includes write-out of event files. The LHEH5 files are written to local SSD storage. The event files are not included in this dataset. Additional run parameters e^+ e^- number of batches batch size rate (main variant) rate (native variant) +0j 20 1,179,648 5.7e9 7.4e9 +1j 20 1,179,648 * 2 1.4e10 3.3e10 +2j 80 1,179,648 / 2 1.3e10 5.4e10 +3j 20 1,179,648 7.7e9 1.6e10 +4j 20 1,179,648 2.2e9 3.1e9 +5j 20 1,179,648 3.4e8 4.2e8 The number of batches has been chosen such that the most important sub processes are likely to be sampled and such that the overall runtime is at least a few seconds. The batch size has been chosen by scanning over factors of two and picking the best-performing batch size with the native variant. The event rate is the number given in the Pepper output. It does not include the closing time of the generated HDF5 file, which can be significant for the two lowest multiplicities. This can be checked by inspecting the corresponding relative and absolute timing results in the `*timing.csv` files included in the dataset.

Pepper v1.1.1工具下正负电子与喷注联合产生过程的性能剖析结果 内容说明 本数据集包含以下GPU加速无权重事件生成运行的结果与输出文件: 1. 一段短时运行以生成Pepper内部计时数据(对应`*timers.csv`文件); 2. 与上述相同的运行流程,但搭配`nvprof`性能分析工具(对应`*.nvprof`文件); 3. 单事件批处理运行,执行命令为`ncu --import-source on --set full`(对应`*.ncu-rep`文件)。 本次模拟的物理过程为质心系能量13 TeV下,喷注数n=0~4的正负电子对产生过程。本数据集附带对应的配置文件`pepper.ini`以及`pepper_cache`目录。 所有结果均对应两种Pepper实现变体:其一为借助Kokkos库调用GPU的「主变体(main variant)」,其二为直接使用CUDA原生接口的「原生变体(native variant)」。 软硬件环境与软件栈 本次性能剖析所用的软硬件环境如下: - GPU驱动:NVIDIA-SMI版本550.100,驱动版本550.100,CUDA版本12.4; - GPU型号:Tesla V100S-PCIE-32GB; - CUDA编译工具:发布版11.6,版本号V11.6.124; - Kokkos库版本:4.3.01; - GCC编译器:(GCC) 11.4.1 20231218(Red Hat 11.4.1-3); - CPU型号:Intel(R) Xeon(R) Silver 4214R @ 2.40GHz。 注意事项:本次事件生成流程包含事件文件写入环节,LHEH5格式文件被写入本地SSD存储。本数据集未包含生成的事件文件。 附加运行参数 各工况的运行参数如下表所示: | 喷注数 | 批次数 | 批大小 | 主变体事件率 | 原生变体事件率 | |--------|--------|--------|--------------|----------------| | +0j | 20 | 1,179,648 | 5.7×10⁹ | 7.4×10⁹ | | +1j | 20 | 2×1,179,648 | 1.4×10¹⁰ | 3.3×10¹⁰ | | +2j | 80 | 1,179,648/2 | 1.3×10¹⁰ | 5.4×10¹⁰ | | +3j | 20 | 1,179,648 | 7.7×10⁹ | 1.6×10¹⁰ | | +4j | 20 | 1,179,648 | 2.2×10⁹ | 3.1×10⁹ | | +5j | 20 | 1,179,648 | 3.4×10⁸ | 4.2×10⁸ | 参数说明:本次实验选取批次数的原则为确保核心子过程得到充分采样,且整体运行时长不少于数秒;批大小则通过遍历2的整数倍因子并选取原生变体下性能最优的参数确定。文中的事件率为Pepper输出的原始数值,未包含生成HDF5文件的收尾耗时——对于喷注数最低的两个工况,该耗时占比较为显著。用户可通过本数据集附带的`*timing.csv`文件查看对应的相对与绝对计时结果以验证该结论。

提供机构:
Zenodo
创建时间:
2024-08-09
二维码
社区交流群
二维码
科研交流群
商业服务