Distributed-memory simulations of turbulent flows on modern GPU systems using an adaptive pencil decomposition library
Distributed-memory simulations of turbulent flows on modern GPU systems using an adaptive pencil decomposition library
复制标题
使用自适应铅笔分解库在现代 GPU 系统上对湍流进行分布式内存模拟
DOI:
--
复制
发表时间:
2022
期刊:
影响因子:
--
通讯作者:
M. Fatica
中科院分区:
文献类型:
--
作者:
J. Romero;P. Costa;M. Fatica
This paper presents a performance analysis of pencil domain decomposition methodologies for three-dimensional Computational Fluid Dynamics (CFD) codes for turbulence simulations, on several large GPU-accelerated clusters. The performance was assessed for the numerical solution of the Navier-Stokes equations in two codes which require the calculation of Fast-Fourier Transforms (FFT): a tri-periodic pseudo-spectral solver for isotropic turbulence, and a finite-difference solver for canonical turbulent flows, where the FFTs are used in its Poisson solver. Both codes use a newly developed transpose library that automatically determines the optimal domain decomposition and communication backend on each system. We compared the performance across systems with very different node topologies and available network bandwidth, to show how these characteristics impact decomposition selection for best performance. Additionally, we assessed the performance of several communication libraries available on these systems, such as Open-MPI, IBM Spectrum MPI, Cray MPI, the NVIDIA Collective Communication Library (NCCL), and NVSHMEM. Our results show that the optimal combination of communication backend and domain decomposition is highly system-dependent, and that the adaptive decomposition library is key in ensuring efficient resource usage with minimal user effort.