Accelerating Restarted GMRES With Mixed Precision Arithmetic

Accelerating Restarted GMRES With Mixed Precision Arithmetic
复制标题

DOI:
10.1109/tpds.2021.3090757
复制
发表时间:
2021-06
影响因子:
5.3
通讯作者:
Neil Lindquist;P. Luszczek;J. Dongarra
Neil Lindquist;P. Luszczek;J. Dongarra
中科院分区:
计算机科学2区
文献类型:
--
作者:
Neil Lindquist;P. Luszczek;J. Dongarra

文献摘要

被引文献

相似文献

广义最小残差法(GMRES)是一种常用的稀疏非对称线性方程组的迭代Krylov求解器。与其他迭代求解器一样,数据移动在其运行时占主导地位。为了提高这种性能,我们建议在降低精度的情况下运行GMRES,而关键操作保持在全精度。此外,我们提供的理论结果连接有限精度GMRES与经典的Gram-Schmidt与再正交化(CGSR)和其无限精度对应的收敛性,这有助于证明该方法的收敛性的双精度。我们在GPU加速节点上使用各种矩阵和预处理器测试了混合精度方法。排除不完全LU分解而不填充(ILU(0))预处理器,我们实现了相对于可比双精度实现的平均加速比,范围从8%到61%,更简单的预处理器实现了更高的加速比。
The generalized minimum residual method (GMRES) is a commonly used iterative Krylov solver for sparse, non-symmetric systems of linear equations. Like other iterative solvers, data movement dominates its run time. To improve this performance, we propose running GMRES in reduced precision with key operations remaining in full precision. Additionally, we provide theoretical results linking the convergence of finite precision GMRES with classical Gram-Schmidt with reorthogonalization (CGSR) and its infinite precision counterpart which helps justify the convergence of this method to double-precision accuracy. We tested the mixed-precision approach with a variety of matrices and preconditioners on a GPU-accelerated node. Excluding the incomplete LU factorization without fill in (ILU(0)) preconditioner, we achieved average speedups ranging from 8 to 61 percent relative to comparable double-precision implementations, with the simpler preconditioners achieving the higher speedups.