Techniques, Tricks, and Algorithms for Efficient GPU-Based Processing of Higher Order Hyperbolic PDEs

Techniques, Tricks, and Algorithms for Efficient GPU-Based Processing of Higher Order Hyperbolic PDEs
复制标题

基于 GPU 高效处理高阶双曲偏微分方程的技术、技巧和算法

DOI:
10.1007/s42967-022-00235-9
复制
发表时间:
2023
影响因子:
1.6
通讯作者:
Kumar, Harish
Kumar, Harish
中科院分区:
数学4区
文献类型:
--
作者:
Subramanian, Sethupathy;Balsara, Dinshaw S.;Bhoriya, Deepak;Kumar, Harish

文献摘要

参考文献

相似文献

预计GPU计算将在所有现代百亿亿级超级计算机中发挥不可或缺的作用。预计高阶Godunov格式将在这种超级计算机的应用组合中占相当大的比例。因此,让双曲偏微分方程的高阶方案的用户群体为这个新出现的机会做好准备是非常重要的。并不是每一种用于双曲偏微分方程解的时空更新的算法都能很好地适用于gpu。然而,我们确定了一小部分算法的核心,这些算法非常适合GPU计算。基于对可用选项的分析,我们已经能够确定用于空间重建的加权本质非振荡(WENO)算法以及用于时间扩展的任意导数(ADER)算法,然后是校正步骤,作为获胜的三部分算法组合。即使已经确定了一个获胜的算法子集,也不清楚它们是否会无缝地移植到gpu上。CPU和GPU之间的低数据吞吐量,以及现代GPU上非常小的缓存大小,意味着我们必须考虑将应用程序移植到GPU任务的所有方面。出于这个原因,本文确定了将这类非常有用的高阶算法成功移植到gpu所需的技术和技巧。应用程序代码面临着进一步的挑战——GPU的结果需要与CPU的结果几乎无法区分——以便在GPU移植期间保留嵌入在这些应用程序代码中的遗留知识库。这个要求常常使完全重写代码变得不可能。出于这个原因,使用基于OpenACC指令的方法是最安全的,这样大部分代码保持完整(只要它最初写得很好)。本文旨在为任何寻求基于openacc的高阶Godunov方案到gpu的端口的人提供一站式服务。我们将重点放在使用高阶Godunov方案的三个广泛和高影响的领域。第一个领域是计算流体动力学。第二种是计算磁流体力学(MHD),它有一个必须模拟保存的对合约束。第三种是计算电动力学(CED),它具有对合约束和极其刚性的源项。总之,这三种不同的高阶Godunov方法的使用,涵盖了许多最重要的应用领域。在这三种情况下,我们展示了算法、技术和技巧的最佳使用,以及OpenACC的使用,在gpu上产生了最高的速度提升。作为奖励,我们发现了一个最显著和最理想的结果:一些高阶方案,每个区域的操作次数更多,在gpu上表现出比低阶方案更好的加速。换句话说,GPU是克服高阶方案的高计算复杂性的最佳策略。还确定了今后改进的若干途径。本文提出了一个使用gpu和相当数量的高端多核cpu的实际应用程序的可伸缩性研究。研究发现,gpu比同等数量的cpu提供了实质性的性能优势,特别是在使用本文设计的所有方法时。
GPU computing is expected to play an integral part in all modern Exascale supercomputers. It is also expected that higher order Godunov schemes will make up about a significant fraction of the application mix on such supercomputers. It is, therefore, very important to prepare the community of users of higher order schemes for hyperbolic PDEs for this emerging opportunity.Not every algorithm that is used in the space-time update of the solution of hyperbolic PDEs will take well to GPUs. However, we identify a small core of algorithms that take exceptionally well to GPU computing. Based on an analysis of available options, we have been able to identify weighted essentially non-oscillatory (WENO) algorithms for spatial reconstruction along with arbitrary derivative (ADER) algorithms for time extension followed by a corrector step as the winning three-part algorithmic combination. Even when a winning subset of algorithms has been identified, it is not clear that they will port seamlessly to GPUs. The low data throughput between CPU and GPU, as well as the very small cache sizes on modern GPUs, implies that we have to think through all aspects of the task of porting an application to GPUs. For that reason, this paper identifies the techniques and tricks needed for making a successful port of this very useful class of higher order algorithms to GPUs.Application codes face a further challenge—the GPU results need to be practically indistinguishable from the CPU results—in order for the legacy knowledge bases embedded in these applications codes to be preserved during the port of GPUs. This requirement often makes a complete code rewrite impossible. For that reason, it is safest to use an approach based on OpenACC directives, so that most of the code remains intact (as long as it was originally well-written). This paper is intended to be a one-stop shop for anyone seeking to make an OpenACC-based port of a higher order Godunov scheme to GPUs.We focus on three broad and high-impact areas where higher order Godunov schemes are used. The first area is computational fluid dynamics (CFD). The second is computational magnetohydrodynamics (MHD) which has an involution constraint that has to be mimetically preserved. The third is computational electrodynamics (CED) which has involution constraints and also extremely stiff source terms. Together, these three diverse uses of higher order Godunov methodology, cover many of the most important applications areas. In all three cases, we show that the optimal use of algorithms, techniques, and tricks, along with the use of OpenACC, yields superlative speedups on GPUs. As a bonus, we find a most remarkable and desirable result: some higher order schemes, with their larger operations count per zone, show better speedup than lower order schemes on GPUs. In other words, the GPU is an optimal stratagem for overcoming the higher computational complexities of higher order schemes. Several avenues for future improvement have also been identified. A scalability study is presented for a real-world application using GPUs and comparable numbers of high-end multicore CPUs. It is found that GPUs offer a substantial performance benefit over comparable number of CPUs, especially when all the methods designed in this paper are used.
DOI: 10.1006/jcph.1997.5705
发表时间: 1981-01-01
影响因子: 4.1
作者:
ROE, PL
通讯作者: ROE, PL
以 3D 方式模拟磁引导风:I. 磁 O 超巨星的等温模拟
DOI: --
发表时间: 2022
影响因子: 4.8
作者:
Sethupathy S.;D. Balsara;A. ud;M. Gagne
通讯作者: M. Gagne
DOI: 10.1016/j.jcp.2008.05.025
发表时间: 2008-09-10
影响因子: 4.1
作者:
Dumbser, Michael;Balsara, Dinshaw S.;Munz, Claus-Dieter
通讯作者: Munz, Claus-Dieter