Distributed-memory simulations of turbulent flows on modern GPU systems using an adaptive pencil decomposition library

Distributed-memory simulations of turbulent flows on modern GPU systems using an adaptive pencil decomposition library
复制标题

使用自适应铅笔分解库在现代 GPU 系统上对湍流进行分布式内存模拟

DOI:
--
复制
发表时间:
2022
期刊:
Platform for Advanced Scientific Computing Conference
影响因子:
--
通讯作者:
M. Fatica
M. Fatica
中科院分区:
--
文献类型:
--
作者:
J. Romero;P. Costa;M. Fatica

文献摘要

被引文献

相似文献

本文提出了一种性能分析的铅笔区域分解方法的三维计算流体动力学(CFD)代码的湍流模拟,在几个大型GPU加速集群。性能进行了评估的数值解的Navier-Stokes方程的两个代码,需要计算的快速傅立叶变换(FFT):一个三周期的伪谱求解器的各向同性湍流,和一个有限差分求解器的典型湍流,其中的FFT是在其泊松求解器。这两个代码都使用了一个新开发的转置库,可以自动确定每个系统上的最佳域分解和通信后端。我们比较了具有非常不同的节点拓扑和可用网络带宽的系统的性能,以展示这些特征如何影响分解选择以获得最佳性能。此外,我们还评估了这些系统上可用的几个通信库的性能,例如Open-MPI、IBM Spectrum MPI、Cray MPI、NVIDIA Collective Communication Library(NCCL)和NVSHMEM。我们的研究结果表明,通信后端和域分解的最佳组合是高度依赖于系统的,自适应分解库是以最小的用户努力确保有效的资源使用的关键。
This paper presents a performance analysis of pencil domain decomposition methodologies for three-dimensional Computational Fluid Dynamics (CFD) codes for turbulence simulations, on several large GPU-accelerated clusters. The performance was assessed for the numerical solution of the Navier-Stokes equations in two codes which require the calculation of Fast-Fourier Transforms (FFT): a tri-periodic pseudo-spectral solver for isotropic turbulence, and a finite-difference solver for canonical turbulent flows, where the FFTs are used in its Poisson solver. Both codes use a newly developed transpose library that automatically determines the optimal domain decomposition and communication backend on each system. We compared the performance across systems with very different node topologies and available network bandwidth, to show how these characteristics impact decomposition selection for best performance. Additionally, we assessed the performance of several communication libraries available on these systems, such as Open-MPI, IBM Spectrum MPI, Cray MPI, the NVIDIA Collective Communication Library (NCCL), and NVSHMEM. Our results show that the optimal combination of communication backend and domain decomposition is highly system-dependent, and that the adaptive decomposition library is key in ensuring efficient resource usage with minimal user effort.