Basker: A Threaded Sparse LU Factorization Utilizing Hierarchical Parallelism and Data Layouts

Basker: A Threaded Sparse LU Factorization Utilizing Hierarchical Parallelism and Data Layouts
复制标题

Basker:利用分层并行性和数据布局的线程稀疏 LU 分解

DOI:
--
复制
发表时间:
2016
期刊:
IEEE International Symposium on Parallel & Distributed Processing, Workshops and Phd Forum
影响因子:
--
通讯作者:
H. Thornquist
H. Thornquist
中科院分区:
--
文献类型:
--
作者:
J. Booth;S. Rajamanickam;H. Thornquist

文献摘要

被引文献

相似文献

可伸缩稀疏LU分解对于电路和电网的高效数值模拟至关重要。在这项工作中,我们提出了一种新的可扩展稀疏直接求解器,称为Basker。Basker引入了一种新的算法来并行化稀疏LU分解的Gilbert-Peierls算法。随着体系结构的发展,需要有层次结构的算法来匹配线程团队、单个线程和向量级并行性中的层次结构。Basker被设计成很好地映射到架构中的这种层次结构。数据布局还需要匹配内存中的多个层次结构。Basker使用稀疏矩阵的二维层次结构,该结构映射到内存体系结构中的层次结构和并行的层次结构。我们使用来自佛罗里达大学稀疏矩阵集合和Xyce电路模拟的电路和电网矩阵,在英特尔SandyBridge和Xeon Phi平台上对Basker进行了性能评估。相对于KLU, Basker在CPU(16核)上实现了5.91倍的几何平均加速,在Xeon Phi(32核)上实现了7.4倍的几何平均加速。对于低填充电路矩阵,Basker在CPU(16核)上的性能比Intel MKL Pardiso (PMKL)高出30倍,在Xeon Phi(32核)上的性能高出7.5倍。此外,Basker在一个具有挑战性的矩阵序列上提供了5.4倍的加速,这些矩阵序列取自实际的Xyce模拟。
Scalable sparse LU factorization is critical for efficient numerical simulation of circuits and electrical power grids. In this work, we present a new scalable sparse direct solver called Basker. Basker introduces a new algorithm to parallelize the Gilbert-Peierls algorithm for sparse LU factorization. As architectures evolve, there exists a need for algorithms that are hierarchical in nature to match the hierarchy in thread teams, individual threads, and vector level parallelism. Basker is designed to map well to this hierarchy in architectures. There is also a need for data layouts to match multiple levels of hierarchy in memory. Basker uses a two-dimensional hierarchical structure of sparse matrices that maps to the hierarchy in the memory architectures and to the hierarchy in parallelism. We present performance evaluations of Basker on the Intel SandyBridge and Xeon Phi platforms using circuit and power grid matrices taken from the University of Florida sparse matrix collection and from Xyce circuit simulations. Basker achieves a geometric mean speedup of 5.91× on CPU (16 cores) and 7.4× on Xeon Phi (32 cores) relative to KLU. Basker outperforms Intel MKL Pardiso (PMKL) by as much as 30× on CPU (16 cores) and 7.5× on Xeon Phi (32 cores) for low fill-in circuit matrices. Furthermore, Basker provides 5.4× speedup on a challenging matrix sequence taken from an actual Xyce simulation.