Acceleration of the Parameterization of Unified Microphysics Across Scales (PUMAS) on the Graphics Processing Unit (GPU) With Directive‐Based Methods

Acceleration of the Parameterization of Unified Microphysics Across Scales (PUMAS) on the Graphics Processing Unit (GPU) With Directive‐Based Methods
复制标题

DOI:
10.1029/2022ms003515
复制
发表时间:
2023-05
影响因子:
6.8
通讯作者:
Jian Sun;J. Dennis;S. Mickelson;B. Vanderwende;A. Gettelman;K. Thayer‐Calder
Jian Sun;J. Dennis;S. Mickelson;B. Vanderwende;A. Gettelman;K. Thayer‐Calder
中科院分区:
地球科学2区
文献类型:
--
作者:
Jian Sun;J. Dennis;S. Mickelson;B. Vanderwende;A. Gettelman;K. Thayer‐Calder

文献摘要

相似文献

云微物理学是气候模式中最耗时的组成部分之一。在这项研究中,我们端口的云微物理参数化的社区大气模型(CAM),被称为跨尺度的统一微物理参数化(PUMAS),从CPU到GPU,以寻求计算加速。基于指令的方法(OpenACC和OpenMP目标卸载)被确定为最适合我们的开发实践,使单个版本的源代码能够在CPU或GPU上运行,并产生更好的可移植性和可维护性。它们的性能首先在PUMAS独立内核中进行检查,只要GPU上有足够的计算负担,基于指令的方法就可以胜过CPU节点。当我们在实际的CAM模拟中在GPU上运行PUMAS时,观察到一致的行为。PUMAS执行时间(包括CPU和GPU之间的数据移动)的3.6倍加速是在粗略水平分辨率下实现的(8个NVIDIA V100 GPU对36个Intel Skylake CPU内核)。在高分辨率(24个NVIDIA V100 GPU对108个Intel Skylake CPU内核)下,这种加速比进一步提高到5.4倍,这突出了GPU支持更大问题大小的事实。这项研究表明,在CAM模拟中使用GPU可以节省显着的计算成本,即使只有一小部分代码启用了GPU。因此,我们鼓励将更多的参数化移植到GPU,以利用其计算优势。
Cloud microphysics is one of the most time‐consuming components in a climate model. In this study, we port the cloud microphysics parameterization in the Community Atmosphere Model (CAM), known as Parameterization of Unified Microphysics Across Scales (PUMAS), from CPU to GPU to seek a computational speedup. The directive‐based methods (OpenACC and OpenMP target offload) are determined as the best fit specifically for our development practices, which enable a single version of source code to run either on the CPU or GPU, and yield a better portability and maintainability. Their performance is first examined in a PUMAS stand‐alone kernel and the directive‐based methods can outperform a CPU node as long as there is enough computational burden on the GPU. A consistent behavior is observed when we run PUMAS on the GPU in a practical CAM simulation. A 3.6× speedup of the PUMAS execution time, including data movement between CPU and GPU, is achieved at a coarse horizontal resolution (8 NVIDIA V100 GPUs against 36 Intel Skylake CPU cores). This speedup further increases up to 5.4× at a high resolution (24 NVIDIA V100 GPUs against 108 Intel Skylake CPU cores), which highlights the fact that GPU favors larger problem size. This study demonstrates that using GPU in a CAM simulation can save noticeable computational costs even with a small portion of code being GPU‐enabled. Therefore, we are encouraged to port more parameterizations to GPU to take advantage of its computational benefit.