Optimising the performance of the spectral/ h p element method with collective linear algebra operations

Optimising the performance of the spectral/ h p element method with collective linear algebra operations
复制标题

通过集体线性代数运算优化谱/ HP 元方法的性能

DOI:
10.1016/j.cma.2016.07.001
复制
发表时间:
2016
影响因子:
7.2
通讯作者:
Moxey D
Moxey D
中科院分区:
工程技术1区
文献类型:
--
作者:
Moxey D

文献摘要

相似文献

随着计算硬件的发展,核心数量的增加意味着内存带宽正在成为获得数值方法峰值性能的决定性因素。高阶有限元方法,如在谱/h p框架Nektar++中实现的那些方法,特别适合于这种环境。与通常利用稀疏存储的低阶方法不同,表示高阶算子的矩阵具有更大的密度和更丰富的结构。在本文中,我们将展示如何利用这些品质来提高节点上的运行时性能,这些节点包括一个典型的高性能计算系统,通过将多个元素上的关键运算符的操作合并到一个单一的内存高效块中。我们研究了不同的策略,以实现最佳性能的多项式阶数和元素类型的范围。由于这些策略都依赖于外部因素,如BLAS实现和感兴趣的几何形状,我们提出了一种技术,自动选择最有效的策略在运行时。
As computing hardware evolves, increasing core counts mean that memory bandwidth is becoming the deciding factor in attaining peak performance of numerical methods. High-order finite element methods, such as those implemented in the spectral/h p framework Nektar++, are particularly well-suited to this environment. Unlike low-order methods that typically utilise sparse storage, matrices representing high-order operators have greater density and richer structure. In this paper, we show how these qualities can be exploited to increase runtime performance on nodes that comprise a typical high-performance computing system, by amalgamating the action of key operators on multiple elements into a single, memory-efficient block. We investigate different strategies for achieving optimal performance across a range of polynomial orders and element types. As these strategies all depend on external factors such as BLAS implementation and the geometry of interest, we present a technique for automatically selecting the most efficient strategy at runtime.