High-order finite-element seismic wave propagation modeling with MPI on a large GPU cluster

High-order finite-element seismic wave propagation modeling with MPI on a large GPU cluster
复制标题

DOI:
10.1016/j.jcp.2010.06.024
复制
发表时间:
2010-10-01
影响因子:
4.1
通讯作者:
Michea, David
Michea, David
中科院分区:
物理与天体物理2区
文献类型:
--
作者:
Komatitsch, Dimitri;Erlebacher, Gordon;Michea, David

文献摘要

被引文献

相似文献

我们实施了一个高阶有限元应用,该应用程序对地震波传播进行数值模拟,例如,在大陆的规模上或在石油行业的主动地震收购实验中,在一大群Nvidia Tesla Tesla图形上产生了地震繁殖。使用CUDA编程环境和基于MPI传递的非阻滞消息的卡片。与许多有限元的实现相反,我们的实现成功地实现了单一的精确度,从而最大程度地提高了当前GPU的性能。我们讨论了代码的实现和优化,并将其与CPU语言中的现有非常优化的实现进行了比较。我们使用网格着色来有效处理非结构化网格上自由度的求和操作,以及非阻止MPI消息,以与整个网络上的通信以及通过PCIE与设备的数据传输与GPU上的计算重叠。我们执行许多数值测试来验证单精度CUDA和MPI实现并评估其准确性。然后,我们分析性能测量结果,并取决于问题的映射到参考CPU群集的映射方式,我们获得了20倍或12倍的加速。 (c)2010 Elsevier Inc.保留所有权利。
We implement a high-order finite-element application, which performs the numerical simulation of seismic wave propagation resulting for instance from earthquakes at the scale of a continent or from active seismic acquisition experiments in the oil industry, on a large cluster of NVIDIA Tesla graphics cards using the CUDA programming environment and non-blocking message passing based on MPI. Contrary to many finite-element implementations, ours is implemented successfully in single precision, maximizing the performance of current generation GPUs. We discuss the implementation and optimization of the code and compare it to an existing very optimized implementation in C language and MPI on a classical cluster of CPU nodes. We use mesh coloring to efficiently handle summation operations over degrees of freedom on an unstructured mesh, and non-blocking MPI messages in order to overlap the communications across the network and the data transfer to and from the device via PCIe with calculations on the GPU. We perform a number of numerical tests to validate the single-precision CUDA and MPI implementation and assess its accuracy. We then analyze performance measurements and depending on how the problem is mapped to the reference CPU cluster, we obtain a speedup of 20x or 12x. (C) 2010 Elsevier Inc. All rights reserved.