Scaling Structured Multigrid to 500K+ Cores through Coarse-Grid Redistribution

Scaling Structured Multigrid to 500K+ Cores through Coarse-Grid Redistribution
复制标题

通过粗网格重新分配将结构化多重网格扩展到 500K 核心

DOI:
10.1137/17m1146440
复制
发表时间:
2018
期刊:
ArXiv
影响因子:
--
通讯作者:
J. Moulton
J. Moulton
中科院分区:
--
文献类型:
--
作者:
†. Andrewreisner;†. LUKEN.OLSON;‡. J.DAVIDMOULTON;J. Moulton

文献摘要

被引文献

相似文献

由偏微分方程的离散化产生的稀疏线性系统的有效解对于许多基于物理的仿真的性能至关重要。多层次方法的算法最优性使其成为高效并行求解器的良好候选者。然而,高性能计算系统的现代架构继续挑战多级求解器的并行可扩展性。虽然代数多重网格方法是强大的解决各种问题,在现代体系结构中的数据局部性和数据移动的成本越来越重要,激励需要仔细利用结构的问题。 强大的逻辑结构变分多重网格方法,如黑盒多重网格(BoxMG),保持整个多重网格层次结构。这避免了间接和增加的粗网格通信成本典型的并行代数多重网格。然而,结构化多重网格的并行可扩展性受到粗网格问题的挑战,其中通信开销占主导地位的计算。本文介绍了一种通过增量凝聚重分布粗网格问题的算法。在预测性能模型的指导下,该算法为结构化多级求解器提供了鲁棒的重分配决策。 一个二维的扩散问题是用来证明显着的增益,在性能上的这种算法比以前的方法,使用凝聚到一个处理器。此外,这种方法的并行可扩展性证明了两个大规模的计算系统,解决高达500K+核心。
The efficient solution of sparse, linear systems resulting from the discretization of partial differential equations is crucial to the performance of many physics-based simulations. The algorithmic optimality of multilevel approaches for common discretizations makes them a good candidate for an efficient parallel solver. Yet, modern architectures for high-performance computing systems continue to challenge the parallel scalability of multilevel solvers. While algebraic multigrid methods are robust for solving a variety of problems, the increasing importance of data locality and cost of data movement in modern architectures motivates the need to carefully exploit structure in the problem. Robust logically structured variational multigrid methods, such as Black Box Multigrid (BoxMG), maintain structure throughout the multigrid hierarchy. This avoids indirection and increased coarse-grid communication costs typical in parallel algebraic multigrid. Nevertheless, the parallel scalability of structured multigrid is challenged by coarse-grid problems where the overhead in communication dominates computation. In this paper, an algorithm is introduced for redistributing coarse-grid problems through incremental agglomeration. Guided by a predictive performance model, this algorithm provides robust redistribution decisions for structured multilevel solvers. A two-dimensional diffusion problem is used to demonstrate the significant gain in performance of this algorithm over the previous approach that used agglomeration to one processor. In addition, the parallel scalability of this approach is demonstrated on two large-scale computing systems, with solves on up to 500K+ cores.