URANOS-2.0: Improved performance, enhanced portability, and model extension towards exascale computing of high-speed engineering flows
收藏资源简介:
We present URANOS-2.0, the second major release of our massively parallel, GPU-accelerated solver for compressible wall flow applications. This latest version represents a significant leap forward in our initial tool, which was launched in 2023 (De Vanna et al. [1]), and has been specifically optimized to take full advantage of the opportunities offered by the cutting-edge pre-exascale architectures available within the EuroHPC JU. In particular, URANOS-2.0 emphasizes portability and compatibility improvements with the two top-ranked supercomputing architectures in Europe: LUMI and Leonardo. These systems utilize different GPU architectures, AMD and NVIDIA, respectively, which necessitates extensive efforts to ensure seamless usability across their distinct structures. In pursuit of this objective, the current release adheres to the OpenACC standard. This choice not only facilitates efficient utilization of the full potential inherent in these extensive GPU-based architectures but also upholds the principles of vendor neutrality, a distinctive characteristic of URANOS solvers in the CFD solvers' panorama. However, the URANOS-2.0 version goes beyond the goals of improving usability and portability; it introduces performance enhancements and restructures the most demanding computational kernels. This translates into a 2× speedup over the same architecture. In addition to its enhanced single-GPU performance, the present solver release demonstrates very good scalability in multi-GPU environments. URANOS-2.0, in fact, achieves strong scaling efficiencies of over 80% across 64 compute nodes (256 GPUs) for both LUMI and Leonardo. Furthermore, its weak scaling efficiencies reach approximately 95% and 90% on LUMI and Leonardo, respectively, when up to 256 nodes (1024 GPUs) are considered. These significant performance advancements position URANOS-2.0 as a state-of-the-art supercomputing platform tailored for compressible wall turbulence applications, establishing the solver as an integrated tool for various aerospace and energy engineering applications, which can span from direct numerical simulations, wall-resolved large eddy simulations, up to most recent wall-modeled large eddy simulations.
本研究推出URANOS-2.0——面向可压缩壁面流动应用的大规模并行GPU加速求解器的第二代正式重大版本。该版本相较于我们2023年发布的初代工具(De Vanna等人[1])实现了跨越式升级,并针对欧洲高性能计算联合企业(EuroHPC JU)旗下前沿的准百亿亿次超级计算架构进行了专项优化,以充分挖掘其性能潜力。具体而言,URANOS-2.0重点优化了可移植性与兼容性,以适配欧洲排名前二的两款超级计算架构:LUMI与Leonardo。这两款架构分别采用AMD与NVIDIA的差异化GPU体系,因此需要开展大量研发工作以确保求解器在二者的不同硬件架构上均可流畅运行。为达成这一目标,本版本严格遵循OpenACC标准。这一选择不仅有助于充分释放这类大规模GPU集群的内在性能潜力,同时坚守了厂商中立的原则——这也是URANOS系列求解器在计算流体动力学(Computational Fluid Dynamics, CFD)求解器领域的鲜明特色。但URANOS-2.0的升级并不止步于易用性与可移植性的提升:它新增了性能优化模块,并重构了计算开销最高的核心计算内核。这使得单架构下的运行速度较前代提升了2倍。除单GPU性能得到增强外,本版本求解器在多GPU环境下同样展现出优异的可扩展性。以LUMI与Leonardo架构为例,URANOS-2.0在64个计算节点(256块GPU)的部署规模下,强缩放效率均超过80%。而当扩展至256个节点(1024块GPU)的规模时,LUMI与Leonardo架构下的弱缩放效率分别可达约95%与90%。上述显著的性能升级使URANOS-2.0成为面向可压缩壁面湍流应用的顶尖超级计算平台,同时将该求解器打造为适用于多类航空航天与能源工程场景的一体化工具,其应用覆盖范围从直接数值模拟(Direct Numerical Simulation, DNS)、壁面解析大涡模拟,直至最新的壁面建模大涡模拟。




