An analysis of the feasibility and benefits of GPU/multicore acceleration of the Weather Research and Forecasting model

An analysis of the feasibility and benefits of GPU/multicore acceleration of the Weather Research and Forecasting model
复制标题

DOI:
10.1002/cpe.3522
复制
发表时间:
2016-05
期刊:
Concurrency and Computation: Practice and Experience
影响因子:
--
通讯作者:
W. Vanderbauwhede;T. Takemi
W. Vanderbauwhede;T. Takemi
中科院分区:
其他
文献类型:
--
作者:
W. Vanderbauwhede;T. Takemi

文献摘要

被引文献

相似文献

人们越来越需要在更短的时间尺度内提供更准确的气候和天气模拟,特别是为了防范飓风和暴雨等恶劣天气事件。由于气候变化,这类事件的严重程度和频率--以及由此产生的经济影响--将急剧上升。使用图形处理单元(GPU)或现场可编程门阵列(FGA)的硬件加速可能会导致运行时间大大缩短或模拟精度更高。本文介绍了天气研究和预报(WRF)模式的研究结果,以评估这种类型的数值天气预报(NWP)程序的GPU和多核加速是否可行和值得。本文的重点是通过将部分代码卸载到诸如GPU之类的加速器来加速在单个计算节点上运行的代码。WRF模式的控制方程组基于多物理过程的可压缩、非静力大气运动。通过讨论它对多物理流体力学程序的更广泛的适用性,我们把这项工作放在了背景中:在许多流体动力学程序中,对流项的数值格式是基于相邻单元之间的有限差分,类似于WRF程序。对于包括多物理过程的流体系统,有许多对这些平流程序的调用。这类数字代码将受益于硬件加速。我们研究了WRF模型的原始代码的性能,提出了一个简单的模型来比较多核CPU和GPU的性能。基于对典型WRF运行的大量廓线分析结果,我们重点研究了标量平流模块的加速。我们讨论了该模块在OpenCL和OpenMP中作为数据并行内核的实现。我们展示了我们的数据并行内核版本的标量平流模块在GPU上的运行速度比在CPU上的原始代码快7倍。然而,由于GPU和CPU之间的数据传输成本非常高(根据我们的分析),完全集成的代码只有很小的加速(两倍)。我们表明,通过GPU加速较大部分的Dynamic代码来抵消数据传输成本是可能的。为了开展这项研究,我们还开发了一个可扩展的软件系统,用于将OpenCL代码集成到WRF等大型Fortran代码库中。这是我们工作的主要贡献之一。我们讨论该系统是为了展示它如何允许在对原始代码进行最少更改(字面上只有几行)的情况下,将原始代码库的部分替换为它们的OpenCL对应部分。我们最后的评估是,即使在目前的系统架构下,以高达五倍的倍数加速WRF-因此也包括其他类似类型的多物理流体动力学代码-绝对是一个可以实现的目标。加速包括NWP程序在内的多物理流体动力学程序对于其在天气预报、环境污染预警和危险物质扩散应急响应中的应用至关重要。实现流体动力学和NWP代码的硬件加速功能是构建最新和未来计算机体系结构的先决条件。版权所有©2015 John Wiley&Sons,Ltd.
There is a growing need for ever more accurate climate and weather simulations to be delivered in shorter timescales, in particular, to guard against severe weather events such as hurricanes and heavy rainfall. Due to climate change, the severity and frequency of such events – and thus the economic impact – are set to rise dramatically. Hardware acceleration using graphics processing units (GPUs) or Field‐Programmable Gate Arrays (FPGAs) could potentially result in much reduced run times or higher accuracy simulations. In this paper, we present the results of a study of the Weather Research and Forecasting (WRF) model undertaken in order to assess if GPU and multicore acceleration of this type of numerical weather prediction (NWP) code is both feasible and worthwhile. The focus of this paper is on acceleration of code running on a single compute node through offloading of parts of the code to an accelerator such as a GPU. The governing equations set of the WRF model is based on the compressible, non‐hydrostatic atmospheric motion with multi‐physics processes. We put this work into context by discussing its more general applicability to multi‐physics fluid dynamics codes: in many fluid dynamics codes, the numerical schemes of the advection terms are based on finite differences between neighboring cells, similar to the WRF code. For fluid systems including multi‐physics processes, there are many calls to these advection routines. This class of numerical codes will benefit from hardware acceleration. We studied the performance of the original code of the WRF model and proposed a simple model for comparing multicore CPU and GPU performance. Based on the results of extensive profiling of representative WRF runs, we focused on the acceleration of the scalar advection module. We discuss the implementation of this module as a data‐parallel kernel in both OpenCL and OpenMP. We show that our data‐parallel kernel version of the scalar advection module runs up to seven times faster on the GPU compared with the original code on the CPU. However, as the data transfer cost between GPU and CPU is very high (as shown by our analysis), there is only a small speed‐up (two times) for the fully integrated code. We show that it would be possible to offset the data transfer cost through GPU acceleration of a larger portion of the dynamics code. In order to carry out this research, we also developed an extensible software system for integrating OpenCL code into large Fortran code bases such as WRF. This is one of the main contributions of our work. We discuss the system to show how it allows the replacement of the sections of the original codebase with their OpenCL counterparts with minimal changes – literally only a few lines – to the original code. Our final assessment is that, even with the current system architectures, accelerating WRF – and hence also other, similar types of multi‐physics fluid dynamics codes – with a factor of up to five times is definitely an achievable goal. Accelerating multi‐physics fluid dynamics codes including NWP codes is vital for its application to weather forecasting, environmental pollution warning, and emergency response to the dispersion of hazardous materials. Implementing hardware acceleration capability for fluid dynamics and NWP codes is a prerequisite for up‐to‐date and future computer architectures. Copyright © 2015 John Wiley & Sons, Ltd.