Parallel preconditioned conjugate gradient algorithm on GPU

Parallel preconditioned conjugate gradient algorithm on GPU
复制标题

DOI:
10.1016/j.cam.2011.04.025
复制
发表时间:
2012-09
期刊:
J. Comput. Appl. Math.
影响因子:
--
通讯作者:
Rudi Helfenstein;J. Koko
Rudi Helfenstein;J. Koko
中科院分区:
其他
文献类型:
--
作者:
Rudi Helfenstein;J. Koko

文献摘要

被引文献

相似文献

我们提出了在 GPU 平台上并行实现预条件共轭梯度算法。预处理矩阵是从 SSOR 预处理器导出的近似逆矩阵。通过稀疏矩阵向量乘法使用,所提出的预处理器非常适合大规模并行 GPU 架构。与共轭梯度算法的 CPU 实现相比,我们的 GPU 预处理共轭梯度实现速度快了 10 倍(最坏情况下快了 8 倍)。
We propose a parallel implementation of the Preconditioned Conjugate Gradient algorithm on a GPU platform. The preconditioning matrix is an approximate inverse derived from the SSOR preconditioner. Used through sparse matrix–vector multiplication, the proposed preconditioner is well suited for the massively parallel GPU architecture. As compared to CPU implementation of the conjugate gradient algorithm, our GPU preconditioned conjugate gradient implementation is up to 10 times faster (8 times faster at worst).