A two-scale approach for efficient on-the-fly operator assembly in massively parallel high performance multigrid codes

A two-scale approach for efficient on-the-fly operator assembly in massively parallel high performance multigrid codes
复制标题

大规模并行高性能多重网格代码中高效即时算子组装的两种尺度方法

DOI:
10.1016/j.apnum.2017.07.006
复制
发表时间:
2016
期刊:
ArXiv
影响因子:
--
通讯作者:
B. Wohlmuth
B. Wohlmuth
中科院分区:
--
文献类型:
--
作者:
S. Bauer;M. Mohr;U. Rüde;J. Weismüller;M. Wittmann;B. Wohlmuth

文献摘要

参考文献

被引文献

相似文献

大规模无矩阵有限元实现节省内存,往往比使用经典稀疏矩阵技术的实现显着更快。它们特别适合于大规模并行几何多重网格求解器结合层次混合网格多面体域。在常数系数的情况下,不同模板条目的数量仅取决于粗网格,并且不随细化级别的数量而增加。然而,对于非多面体域,情况发生了变化。然后,即使对于拉普拉斯算子,元素映射也会导致精细的网格扩展,该网格扩展可以从网格点到网格点而变化。传统的无矩阵技术是基于一个元素的组装,然后导致在计算成本的显着增加。为了弥补这一缺点,我们引入了一种使用代理运算符的新的双尺度方法。它利用了一个分段多项式近似的细网格运营商的模板的条目相对于粗网格尺寸。这些替代多项式的低成本评估结果在一个有效的模板组装非多面体域的飞行。我们讨论和说明数字两尺度先验界。如果与双重离散技术相结合,近似解的精度可以进一步提高。仔细的性能分析与基于执行-缓存-内存模型的硬件感知代码优化相结合,可以显著提高速度。弱标度和强标度结果说明了这种新的双尺度方法在大规模PDE模拟中的潜力。
Large scale matrix-free finite element implementations save memory and are often significantly faster than implementations using classical sparse matrix techniques. They are especially well suited for massively parallel geometric multigrid solvers in combination with hierarchical hybrid grids on polyhedral domains. In the case of constant coefficients, the number of different stencil entries depends only on the coarse grid and does not increase with the number of refinement levels. However, for non-polyhedral domains the situation changes. Then even for the Laplace operator, the element mapping leads to fine grid stencils that can vary from grid point to grid point. Traditional matrix-free techniques that are based on an element-wise assembly then result in a considerably increase in computational cost. To compensate for this shortcoming, we introduce a new two-scale approach that uses a surrogate operator. It exploits a piecewise polynomial approximation of the entries of the stencil of the fine grid operator with respect to the coarse mesh size. The low-cost evaluation of these surrogate polynomials results in an efficient stencil assembly on-the-fly for non-polyhedral domains. We discuss and illustrate numerically two-scale a priori bounds. The accuracy of the approximate solution can be further improved if combined with a double discretization technique. A careful performance analysis in combination with a hardware–aware code optimization based on the Execution–Cache–Memory model yields a significant speed up. Weak and strong scaling results illustrate the potential of this new two-scale approach within large scale PDE simulations.
EXA-DUNE:灵活的 PDE 求解器、数值方法和应用
DOI: 10.1007/978-3-319-14313-2_45
发表时间: 2014
期刊:
影响因子: --
作者:
Bastian;Engwer;Göddeke;Ippisch;Ohlberger;Fahlke;Kaulmann;Steffen Müthing
通讯作者: Steffen Müthing