Finite element assembly strategies on multi-core and many-core architectures

Finite element assembly strategies on multi-core and many-core architectures
复制标题

DOI:
10.1002/fld.3648
复制
发表时间:
2013-01-10
影响因子:
1.8
通讯作者:
Sherwin, S. J.
Sherwin, S. J.
中科院分区:
工程技术4区
文献类型:
--
作者:
Markall, G. R.;Slemmer, A.;Sherwin, S. J.

文献摘要

被引文献

相似文献

我们论证了如果要实现它们各自的性能潜力,在多核(CPU)和多核(GPU)架构上需要完全不同的有限元方法(FEM)实现。我们使用有限元推进扩散求解器进行的数值研究表明,只有致力于跨越实现的高级结构的特定和多样化的算法选择,才能提高每个体系结构的性能。为实现单个体系结构的高性能而做出这些承诺会导致性能可移植性的损失。包含冗余数据但支持合并内存访问的数据结构在多核体系结构上速度更快,而间接访问的无冗余数据结构在多核体系结构上速度更快。全局汇编的Addto算法在多核体系结构上是最优的,而局部矩阵方法在多核体系结构上是最优的,尽管需要比Addto算法更多的计算。这些结果表明,在现代高性能体系结构上实现FEMS、谱元素方法和低阶间断Galerkin方法时,正确选择算法和数据结构是有价值的。版权所有(C)2012 John Wiley&Sons,Ltd.
We demonstrate that radically differing implementations of finite element methods (FEMs) are needed on multi-core (CPU) and many-core (GPU) architectures, if their respective performance potential is to be realised. Our numerical investigations using a finite element advectiondiffusion solver show that increased performance on each architecture can only be achieved by committing to specific and diverse algorithmic choices that cut across the high-level structure of the implementation. Making these commitments to achieve high performance for a single architecture leads to a loss of performance portability. Data structures that include redundant data but enable coalesced memory accesses are faster on many-core architectures, whereas redundancy-free data structures that are accessed indirectly are faster on multi-core architectures. The Addto algorithm for global assembly is optimal on multi-core architectures, whereas the Local Matrix Approach is optimal on many-core architectures despite requiring more computation than the Addto algorithm. These results demonstrate the value in making the correct choice of algorithm and data structure when implementing FEMs, spectral element methods and low-order discontinuous Galerkin methods on modern high-performance architectures. Copyright (c) 2012 John Wiley & Sons, Ltd.