The fast multipole method on parallel clusters, multicore processors, and graphics processing units

The fast multipole method on parallel clusters, multicore processors, and graphics processing units
复制标题

DOI:
10.1016/j.crme.2010.12.005
复制
发表时间:
2011-02
影响因子:
0.8
通讯作者:
Eric F Darve;C. Cecka;Toru Takahashi
Eric F Darve;C. Cecka;Toru Takahashi
中科院分区:
工程技术4区
文献类型:
--
作者:
Eric F Darve;C. Cecka;Toru Takahashi

文献摘要

被引文献

相似文献

在本文中,我们讨论如何在从计算机集群到多核处理器和显卡(GPU)的现代并行计算机上实现快速多极子方法(FMM)。FMM对于并行计算来说是一个有点困难的应用程序,因为它是树形结构,而且它需要许多复杂的操作,而这些操作并不是规则结构的。例如,具有密集矩阵的计算线性代数允许许多利用规则计算模式的优化。FMM也可以进行类似的优化,但我们将看到优化步骤的复杂性更大。讨论将从FMMS的一般介绍开始。我们简要讨论了FMM的并行方法,如并行构建FMM树,减少FMM过程中的通信。最后,我们将重点介绍如何在GPU上移植和优化FMM。
In this article, we discuss how the fast multipole method (FMM) can be implemented on modern parallel computers, ranging from computer clusters to multicore processors and graphics cards (GPU). The FMM is a somewhat difficult application for parallel computing because of its tree structure and the fact that it requires many complex operations which are not regularly structured. Computational linear algebra with dense matrices for example allows many optimizations that leverage the regular computation pattern. FMM can be similarly optimized but we will see that the complexity of the optimization steps is greater. The discussion will start with a general presentation of FMMs. We briefly discuss parallel methods for the FMM, such as building the FMM tree in parallel, and reducing communication during the FMM procedure. Finally, we will focus on porting and optimizing the FMM on GPUs.