New Scheduling Strategies and Hybrid Programming for a Parallel Right-looking Sparse LU Factorization Algorithm on Multicore Cluster Systems

New Scheduling Strategies and Hybrid Programming for a Parallel Right-looking Sparse LU Factorization Algorithm on Multicore Cluster Systems
复制标题

多核集群系统上并行右视稀疏 LU 分解算法的新调度策略和混合编程

DOI:
10.1109/ipdps.2012.63
复制
发表时间:
2012
期刊:
2012 IEEE 26th International Parallel and Distributed Processing Symposium
影响因子:
--
通讯作者:
X. Li
X. Li
中科院分区:
--
文献类型:
--
作者:
I. Yamazaki;X. Li

文献摘要

参考文献

被引文献

相似文献

并行稀疏LU分解是求解大型线性方程组的关键计算核心。在本文中,我们提出了两种策略来解决现代HPC系统上的因子分解算法的一些可扩展性问题。第一种策略是在算法级,尽可能快地调度独立任务,以减少算法的空闲时间和关键路径。我们证明使用数千个核心,我们的新调度策略减少了近三倍的运行时间,从一个国家的最先进的流水线因式分解算法。第二个策略是在编程和架构级别,我们将轻量级的开放MP线程在每个MPI进程中,以减少内存和时间开销的纯MPI实现在许多核心NUMA架构。使用这种混合编程范式,我们获得了显着减少内存使用,同时实现了一个纯MPI范式的并行效率竞争。因此,与由于每核存储器约束而失败的纯MPI范例相比,混合范例可以在每个节点上利用更多的核,并减少相同数量的节点上的因子分解时间。我们展示了广泛的性能分析的新策略,使用数以千计的核心的两个领先的HPC系统,一个Cray-XE 6和IBM iDataBase。
Parallel sparse LU factorization is a key computational kernel in the solution of a large-scale linear system of equations. In this paper, we propose two strategies to address some scalability issues of a factorization algorithm on modern HPC systems. The first strategy is at the algorithmic-level, we schedule independent tasks as soon as possible to reduce the idle time and the critical path of the algorithm. We demonstrate using thousands of cores that our new scheduling strategy reduces the runtime by nearly three-fold from that of a state-of-the-art pipelined factorization algorithm. The second strategy is at both programming- and architecture-levels, we incorporate light-weight Open MP threads in each MPI process to reduce both memory and time overheads of a pure MPI implementation on many core NUMA architectures. Using this hybrid programming paradigm, we obtain a significant reduction in memory usage while achieving a parallel efficiency competitive with that of a pure MPI paradigm. As a result, in comparison to a pure MPI paradigm which failed due to the per-core memory constraint, the hybrid paradigm could utilize more cores on each node and reduce the factorization time on the same number of nodes. We show extensive performance analysis of the new strategies using thousands of cores of the two leading HPC systems, a Cray-XE6 and an IBM iDataPlex.
DOI: 10.1177/1094342010391989
发表时间: 2011-02-01
影响因子: 3.1
作者:
Dongarra, Jack;Beckman, Pete;Yelick, Kathy
通讯作者: Yelick, Kathy