3D finite difference computation on GPUs using CUDA

3D finite difference computation on GPUs using CUDA
复制标题

DOI:
10.1145/1513895.1513905
复制
发表时间:
2009-03
影响因子:
3
通讯作者:
P. Micikevicius
P. Micikevicius
中科院分区:
生物学4区
文献类型:
--
作者:
P. Micikevicius

文献摘要

被引文献

相似文献

在本文中,我们描述了一个GPU并行化的三维有限差分计算使用CUDA。数据访问冗余被用作度量来确定仅模板计算以及波动方程的离散化的最佳实现,这是目前在地震计算中非常感兴趣的。对于更大的扩展,所描述的方法在单个Tesla 10系列GPU上实现了每秒2,400到超过3,000百万个输出点的吞吐量。这大约比运行地震行业类似代码的4核Harpertown CPU高一个数量级。还描述了多GPU并行化,通过将GPU间通信与计算重叠来实现GPU的线性缩放。
In this paper we describe a GPU parallelization of the 3D finite difference computation using CUDA. Data access redundancy is used as the metric to determine the optimal implementation for both the stencil-only computation, as well as the discretization of the wave equation, which is currently of great interest in seismic computing. For the larger stencils, the described approach achieves the throughput of between 2,400 to over 3,000 million of output points per second on a single Tesla 10-series GPU. This is roughly an order of magnitude higher than a 4-core Harpertown CPU running a similar code from seismic industry. Multi-GPU parallelization is also described, achieving linear scaling with GPUs by overlapping inter-GPU communication with computation.