A GPU-accelerated continuous and discontinuous Galerkin non-hydrostatic atmospheric model

A GPU-accelerated continuous and discontinuous Galerkin non-hydrostatic atmospheric model
复制标题

GPU 加速的连续和不连续伽辽金非静水大气模型

DOI:
10.1177/1094342017694427
复制
发表时间:
2019
期刊:
The International Journal of High Performance Computing Applications
影响因子:
--
通讯作者:
F. Giraldo
F. Giraldo
中科院分区:
--
文献类型:
--
作者:
D. Abdi;L. Wilcox;T. Warburton;F. Giraldo

文献摘要

被引文献

相似文献

我们提出了一种图形处理单元(GPU)加速的节点不连续Galerkin方法,用于求解控制大气运动和热力学状态的三维欧拉方程。大气模式动力核心的加速不仅对于更快地获得每日预报,而且对于在给定的模拟时限内获得更准确(高分辨率)的结果,具有重要的实际意义。我们使用适合于GPU的单指令多线程体系结构的算法,使我们的模型相对于CPU的一个核心提高了两个数量级。在Titan超级计算机的一个节点上的测试表明,与16核AMD皓龙CPU相比,使用K20X GPU的加速比高达15倍。使用16,384个 图形处理器对多图形处理器实现的可扩展性进行了测试,结果显示弱扩展效率约为90%。最后,使用几个代表不同尺度大气动力学的基准问题验证了我们的GPU实现的准确性和性能。
We present a Graphics Processing Unit (GPU)-accelerated nodal discontinuous Galerkin method for the solution of the three-dimensional Euler equations that govern the motion and thermodynamic state of the atmosphere. Acceleration of the dynamical core of atmospheric models plays an important practical role in not only getting daily forecasts faster, but also in obtaining more accurate (high resolution) results within a given simulation time limit. We use algorithms suitable for the single instruction multiple thread architecture of GPUs to accelerate our model by two orders of magnitude relative to one core of a CPU. Tests on one node of the Titan supercomputer show a speedup of up to 15 times using the K20X GPU as compared to that on the 16-core AMD Opteron CPU. The scalability of the multi-GPU implementation is tested using 16,384 GPUs, which resulted in a weak scaling efficiency of about 90%. Finally, the accuracy and performance of our GPU implementation is verified using several benchmark problems representative of different scales of atmospheric dynamics.