A new data conversion method for mixed precision Krylov solvers with FP16/BF16 Jacobi preconditioners

A new data conversion method for mixed precision Krylov solvers with FP16/BF16 Jacobi preconditioners
复制标题

具有 FP16/BF16 Jacobi 预处理器的混合精度 Krylov 求解器的新数据转换方法

DOI:
10.1145/3578178.3578222
复制
发表时间:
2023
期刊:
HPC Asia 2023
影响因子:
--
通讯作者:
Onodera Naoyuki
Onodera Naoyuki
中科院分区:
--
文献类型:
--
作者:
Ina Takuya;Idomura Yasuhiro;Imamura Toshiyuki;Onodera Naoyuki

文献摘要

参考文献

相似文献

采用Jacobi预调节器的混合精度Krylov解在FP16和BF16等低精度条件下计算Jacobi预调节器时,往往会出现明显的收敛性退化。发现这种收敛性退化是由于数据转换中的舍入误差导致对角优势的丧失。为了解决这一问题,我们提出了一种新的数据转换方法,该方法旨在保持原始矩阵数据的对角优势。通过在NVIDIA V100 gpu上使用FP16/BF16 Jacobi预调节器计算泊松方程、一般最小残差法和双共轭梯度稳定法对所提方法进行了验证。在这里,新的数据转换是通过切换CUDA中最接近四舍五入、向上四舍五入和向零四舍五入的内在特性来实现的,并且在主迭代之前被调用一次。因此,新数据转换的成本可以忽略不计。当矩阵的系数通过线性系统的缩放而连续变化时,基于最接近四舍五入的传统数据转换根据对角系数和非对角系数的舍入误差的差异呈现出收敛性的周期性变化。这里,收敛退化的周期和幅度取决于有效位长度。另一方面,所提出的数据转换方法完全避免了收敛退化,并且在没有额外开销的情况下实现了Jacobi预调节器的鲁棒混合精度计算。
Mixed precision Krylov solvers with the Jacobi preconditioner often show significant convergence degradation when the Jacobi preconditioner is computed in low precision such as FP16 and BF16. It is found that this convergence degradation is attributed to loss of diagonal dominance due to roundoff errors in data conversion. To resolve this issue, we propose a new data conversion method, which is designed to keep diagonal dominance of the original matrix data. The proposed method is tested by computing the Poisson equation using the conjugate gradient method, the general minimum residual method, and the biconjugate gradient stabilized method with the FP16/BF16 Jacobi preconditioner on NVIDIA V100 GPUs. Here, the new data conversion is implemented by switching the round-nearest, round-up, round-down, and round-towards-zero intrinsics in CUDA, and is called once before the main iteration. Therefore, the cost of the new data conversion is negligible. When the coefficients of matrix is continuously changed by scaling the linear system, the conventional data conversion based on the round-nearest intrinsic shows periodic changes of the convergence property depending on the difference of the roundoff errors between diagonal and off-diagonal coefficients. Here, the period and magnitude of the convergence degradation depend on the bit length of significand. On the other hand, the proposed data conversion method is shown to fully avoid the convergence degradation, and robust mixed precision computing is enabled for the Jacobi preconditioner without extra overheads.
使用混合精度避免通信的 Krylov 方法加速聚变等离子体湍流模拟
DOI: --
发表时间: 2020
期刊:
影响因子: --
作者:
Y. Idomura;T. Ina;Y. Ali;T. Imamura
通讯作者: T. Imamura
2020年IEEE/ACM第11届大规模系统可扩展算法最新进展研讨会(ScalA)
DOI: 10.1109/scala51936.2020
发表时间: 2020
期刊: Concurrency and Computation: Practice and Experience
影响因子: --
作者:
Md Irteja Islam;Verity L Chadwick;A. Martiniuk
通讯作者: A. Martiniuk
DOI: 10.1137/1.9780898718003
发表时间: 2003-05
期刊: --
影响因子: --
作者:
Y. Saad
通讯作者: Y. Saad
Fugaku 上 EFlop/s HPL-AI 基准的实现和数值技术
DOI: 10.1109/scala51936.2020.00014
发表时间: 2020
期刊: 2020 IEEE/ACM 11th Workshop on Latest Advances in Scalable Algorithms for Large-Scale Systems (ScalA)
影响因子: --
作者:
Shuhei Kudo;Keigo Nitadori;Takuya Ina;Toshiyuki Imamura
通讯作者: Toshiyuki Imamura