Optimizing Sparse Matrix-Multiple Vectors Multiplication for Nuclear Configuration Interaction Calculations

Optimizing Sparse Matrix-Multiple Vectors Multiplication for Nuclear Configuration Interaction Calculations
复制标题

DOI:
10.1109/ipdps.2014.125
复制
发表时间:
2014-05
期刊:
2014 IEEE 28th International Parallel and Distributed Processing Symposium
影响因子:
--
通讯作者:
H. Aktulga;A. Buluç;Samuel Williams;Chao Yang
H. Aktulga;A. Buluç;Samuel Williams;Chao Yang
中科院分区:
其他
文献类型:
--
作者:
H. Aktulga;A. Buluç;Samuel Williams;Chao Yang

文献摘要

被引文献

相似文献

利用组态相互作用(CI)方法对轻原子核的性质进行高精度预测需要计算多体核哈密顿矩阵的少数极值本征对。在核的多体费米子动力学(MFDn)代码中,块本征解算器用于此目的。由于所涉及的稀疏矩阵的大尺寸,在本征值计算上花费的时间的显著部分与稀疏矩阵(以及该矩阵的转置)与多个向量(SpMM和SpMM_T)的乘法相关联。SpMM和SpMM_T的现有实现明显低于预期。因此,在本文中,我们提出并分析优化的实现SpMM和SpMM_T。我们的实现基于压缩稀疏块(CSB)矩阵格式和多核架构的目标系统。我们开发了一个性能模型,使我们能够理解和估计我们的SpMM内核实现的性能特征,并证明了我们的实现的效率上一系列的现实世界中的矩阵提取MFDn。特别是,我们获得了3-4的加速比良好的实现的基础上常用的压缩稀疏行(CSR)矩阵格式的必要操作。SpMM内核的改进表明,我们可以在MFDn中使用的块特征求解器的总体执行时间上获得大约40%的速度。
Obtaining highly accurate predictions on the properties of light atomic nuclei using the configuration interaction (CI) approach requires computing a few extremal Eigen pairs of the many-body nuclear Hamiltonian matrix. In the Many-body Fermion Dynamics for nuclei (MFDn) code, a block Eigen solver is used for this purpose. Due to the large size of the sparse matrices involved, a significant fraction of the time spent on the Eigen value computations is associated with the multiplication of a sparse matrix (and the transpose of that matrix) with multiple vectors (SpMM and SpMM_T). Existing implementations of SpMM and SpMM_T significantly underperform expectations. Thus, in this paper, we present and analyze optimized implementations of SpMM and SpMM_T. We base our implementation on the compressed sparse blocks (CSB) matrix format and target systems with multi-core architectures. We develop a performance model that allows us to understand and estimate the performance characteristics of our SpMM kernel implementations, and demonstrate the efficiency of our implementation on a series of real-world matrices extracted from MFDn. In particular, we obtain 3-4 speedup on the requisite operations over good implementations based on the commonly used compressed sparse row (CSR) matrix format. The improvements in the SpMM kernel suggest we may attain roughly a 40% speed up in the overall execution time of the block Eigen solver used in MFDn.