High-performance algebraic multigrid solver optimized for multi-core based distributed parallel systems

High-performance algebraic multigrid solver optimized for multi-core based distributed parallel systems
复制标题

针对基于多核的分布式并行系统优化的高性能代数多重网格求解器

DOI:
10.1145/2807591.2807603
复制
发表时间:
2015
期刊:
SC15: International Conference for High Performance Computing, Networking, Storage and Analysis
影响因子:
--
通讯作者:
P. Dubey
P. Dubey
中科院分区:
--
文献类型:
--
作者:
Jongsoo Park;M. Smelyanskiy;U. Yang;Dheevatsa Mudigere;P. Dubey

文献摘要

被引文献

相似文献

代数多机(AMG)是一种线性求解器,其线性计算复杂性和出色的并行性可伸缩性,预计AMG将成为能够传递数百个Pflops的新兴极限系统的选择AMG的水平性能通常受内存带宽的限制,由于高度稀疏的不规则计算,例如三重稀疏基质产物,稀疏的matrix密集载体乘法,独立集合式浓缩算法以及例如高斯 - 高斯 - 高斯 - 高斯 - 高斯 - SEIDEL。与Hypre基线实现相比,我们开发并分析了高度优化的AMG实现。 - 主题优化,当多个节点较小时,这将转化为类似的高速加速,此外,与NVIDIA(在K40C上运行的AMG)相比,我们的实现达到了1.3倍的速度。
Algebraic Multigrid (AMG) is a linear solver, well known for its linear computational complexity and excellent parallelization scalability. As a result, AMG is expected to be a solver of choice for emerging extreme scale systems capable of delivering hundred Pflops and beyond. While node level performance of AMG is generally limited by memory bandwidth, achieving high bandwidth efficiency is challenging due to highly sparse irregular computation, such as triple sparse matrix products, sparse-matrix dense-vector multiplications, independent set coarsening algorithms, and smoothers such as Gauss-Seidel. We develop and analyze a highly optimized AMG implementation, based on the well-known HYPRE library. Compared to the HYPRE baseline implementation, our optimized implementation achieves 2.0x speedup on a recent Intel® Xeon® Haswell processor. Combined with our other multi-node optimizations, this translates into similarly high speedups when weak-scaled multiple nodes. In addition, our implementation achieves 1.3x speedup compared to AmgX, NVIDIA's high-performance implementation of AMG, running on K40c.