GPU acceleration of a petascale application for turbulent mixing at high Schmidt number using OpenMP 4.5

GPU acceleration of a petascale application for turbulent mixing at high Schmidt number using OpenMP 4.5
复制标题

DOI:
10.1016/j.cpc.2018.02.020
复制
发表时间:
2018-07
期刊:
Comput. Phys. Commun.
影响因子:
--
通讯作者:
M. Clay;M. Clay;D. Buaria;Pui-Kuen Yeung;Toshiyuki Gotoh
M. Clay;M. Clay;D. Buaria;Pui-Kuen Yeung;Toshiyuki Gotoh
中科院分区:
其他
文献类型:
--
作者:
M. Clay;M. Clay;D. Buaria;Pui-Kuen Yeung;Toshiyuki Gotoh

文献摘要

相似文献

本文报告了成功实现大规模并行 GPU 加速算法,用于高施密特数湍流混合的直接数值模拟。这项工作源于最近的一项开发(Comput.Phys.Commun.,第 219 卷,2017 年,313-328),其中显示,当通过专用通信线程重叠通信和计算时,低通信算法可以在 Cray XE6 架构上实现高度可扩展性。现在,在 Cray XK7 架构上使用 OpenMP 4.5 已经实现了更高水平的性能,其中每个节点上 AMD Interlagos 处理器的 16 个整数核心共享一个 Nvidia K20X GPU 加速器。在新算法中,通过在 GPU 上以组合紧凑有限差分 (CCD) 运算的形式执行几乎所有密集标量场计算,可以最大程度地减少数据移动。我们发现,与通常做法不同的内存布局可以为应用 CCD 方案所需的特定内核提供更好的性能。通过将 OpenMP 4.5 NOWAIT 子句添加到 TARGET 构造来启用异步执行,当用于重叠 GPU 上的计算与 CPU 上的计算和通信时,可以提高可扩展性。在美国橡树岭国家实验室的 27 petaflops 超级计算机 Titan 上,对于使用 8192 个 XK7 节点计算的标量场,在 819 2 3 个网格点的最大问题规模下,始终观察到大约 5 倍的 GPU 到 CPU 加速因子。
This paper reports on the successful implementation of a massively parallel GPU-accelerated algorithm for the direct numerical simulation of turbulent mixing at high Schmidt number. The work stems from a recent development (Comput. Phys. Commun., vol. 219, 2017, 313–328), in which a low-communication algorithm was shown to attain high degrees of scalability on the Cray XE6 architecture when overlapping communication and computation via dedicated communication threads. An even higher level of performance has now been achieved using OpenMP 4.5 on the Cray XK7 architecture, where on each node the 16 integer cores of an AMD Interlagos processor share a single Nvidia K20X GPU accelerator. In the new algorithm, data movements are minimized by performing virtually all of the intensive scalar field computations in the form of combined compact finite difference (CCD) operations on the GPUs. A memory layout in departure from usual practices is found to provide much better performance for a specific kernel required to apply the CCD scheme. Asynchronous execution enabled by adding the OpenMP 4.5 NOWAIT clause to TARGET constructs improves scalability when used to overlap computation on the GPUs with computation and communication on the CPUs. On the 27-petaflops supercomputer Titan at Oak Ridge National Laboratory, USA, a GPU-to-CPU speedup factor of approximately 5 is consistently observed at the largest problem size of 819 2 3 grid points for the scalar field computed with 8192 XK7 nodes.