Fast Matrix-Free Evaluation of Discontinuous Galerkin Finite Element Operators

Fast Matrix-Free Evaluation of Discontinuous Galerkin Finite Element Operators
复制标题

DOI:
10.1145/3325864
复制
发表时间:
2017-11
期刊:
ACM Transactions on Mathematical Software (TOMS)
影响因子:
--
通讯作者:
M. Kronbichler;K. Kormann
M. Kronbichler;K. Kormann
中科院分区:
其他
文献类型:
--
作者:
M. Kronbichler;K. Kormann

文献摘要

被引文献

相似文献

我们提出了一种用于不连续伽辽金有限元算子的无矩阵评估的算法框架。它依赖于四边形和六面体网格上的快速求和和分解,针对线性和非线性偏微分方程的一般弱形式。在深入的性能分析中比较不同的算法和数据结构。通过对多个单元和面进行矢量化以及一维插值的偶奇分解来优化局部积分的实现。当从高速缓存运行时,Intel Haswell、Broadwell 和 Knights Landing 处理器可达到高达 60% 的算术峰值,当还考虑从主内存访问向量时,可达到高达 40% 的峰值。在 2×14 Broadwell 核心上,3D 拉普拉斯算子的吞吐量高达每秒 22 亿个未知数,仿射几何上的 3D 平流每秒高达 40 亿个未知数,接近每秒 47 亿个未知数的简单复制操作。我们的实验表明,MPI 幽灵交换对性能有相当大的影响,我们提出了减轻这种影响的策略。最后,讨论了评估几何术语及其性能的各种选项。我们的实现可通过 deal.II 有限元库公开获得。
We present an algorithmic framework for matrix-free evaluation of discontinuous Galerkin finite element operators. It relies on fast quadrature with sum factorization on quadrilateral and hexahedral meshes, targeting general weak forms of linear and nonlinear partial differential equations. Different algorithms and data structures are compared in an in-depth performance analysis. The implementations of the local integrals are optimized by vectorization over several cells and faces and an even-odd decomposition of the one-dimensional interpolations. Up to 60% of the arithmetic peak on Intel Haswell, Broadwell, and Knights Landing processors is reached when running from caches and up to 40% of peak when also considering the access to vectors from main memory. On 2×14 Broadwell cores, the throughput is up to 2.2 billion unknowns per second for the 3D Laplacian and up to 4 billion unknowns per second for the 3D advection on affine geometries, close to a simple copy operation at 4.7 billion unknowns per second. Our experiments show that MPI ghost exchange has a considerable impact on performance and we present strategies to mitigate this effect. Finally, various options for evaluating geometry terms and their performance are discussed. Our implementations are publicly available through the deal.II finite element library.