Conjugate Gradient Solvers with High Accuracy and Bit-wise Reproducibility between CPU and GPU using Ozaki scheme

Conjugate Gradient Solvers with High Accuracy and Bit-wise Reproducibility between CPU and GPU using Ozaki scheme
复制标题

使用 Ozaki 方案在 CPU 和 GPU 之间实现高精度和按位可重复性的共轭梯度求解器

DOI:
10.1145/3432261.3432270
复制
发表时间:
2021
期刊:
Proc. The International Conference on High Performance Computing in Asia-Pacific Region (HPCAsia 2021)
影响因子:
--
通讯作者:
Iakymchuk Roman
Iakymchuk Roman
中科院分区:
--
文献类型:
--
作者:
Mukunoki Daichi;Ozaki Katsuhisa;Ogita Takeshi;Iakymchuk Roman

文献摘要

相似文献

在共轭梯度(CG)等Krylov子空间方法中,由于浮点计算中的舍入误差导致计算精度的损失,迭代次数可能会增加。同时,由于并行计算的计算顺序是不确定的,在不同的计算环境下,即使对于相同的输入,其结果和收敛行为也可能是不相同的。在本研究中,我们在x86 cpu和NVIDIA gpu上提出了一种精确且可重复的非预置CG方法的实现。在我们的方法中,虽然所有变量都存储在FP64上,但所有内部乘积操作(包括矩阵向量乘法)都使用Ozaki方案执行。该方案提供了正确的舍入计算以及不同计算环境之间的位级再现性。在本文中,我们展示了一些例子,其中标准FP64实现的CG在不同的cpu和gpu上导致不相同的结果。然后,我们展示了我们的方法在准确性和可重复性方面的适用性和有效性,以及它们在cpu和gpu上的性能。此外,我们将我们的方法与基于cpu上的精确基本线性代数子程序(ExBLAS)的现有精确和可重复的CG实现的性能进行了比较。
On Krylov subspace methods such as the Conjugate Gradient (CG) method, the number of iterations until convergence may increase due to the loss of computational accuracy caused by rounding errors in floating-point computations. At the same time, because the order of the computation is nondeterministic on parallel computation, the result and the behavior of the convergence may be nonidentical in different computational environments, even for the same input. In this study, we present an accurate and reproducible implementation of the unpreconditioned CG method on x86 CPUs and NVIDIA GPUs. In our method, while all variables are stored on FP64, all inner product operations (including matrix-vector multiplications) are performed using the Ozaki scheme. The scheme delivers the correctly rounded computation as well as bit-level reproducibility among different computational environments. In this paper, we show some examples where the standard FP64 implementation of CG results in nonidentical results across different CPUs and GPUs. We then demonstrate the applicability and the effectiveness of our approach in terms of accuracy and reproducibility and their performance on both CPUs and GPUs. Furthermore, we compare the performance of our method against an existing accurate and reproducible CG implementation based on the Exact Basic Linear Algebra Subprograms (ExBLAS) on CPUs.