An Improved Implementation of Preconditioned Conjugate Gradient Method on GPU

An Improved Implementation of Preconditioned Conjugate Gradient Method on GPU
复制标题

DOI:
10.4304/jsw.7.12.2695-2702
复制
发表时间:
2012-01
期刊:
J. Softw.
影响因子:
--
通讯作者:
Yechen Gui;Guijuan Zhang
Yechen Gui;Guijuan Zhang
中科院分区:
其他
文献类型:
--
作者:
Yechen Gui;Guijuan Zhang

文献摘要

被引文献

相似文献

提出了一种基于CUDA(Compute Unified Device Architecture)的预条件共轭梯度法在GPU上的改进实现。它的目的是解决泊松方程的液体动画与高效率。首先,提出了一种新的存储格式mDIA(modified diagonal storage format),以提高稀疏矩阵向量积(Sparse Matrix-Vector product,SpMV)运算的效率。第二,当使用不完全Cholesky预条件子来探索固有的并行性时,提出了并行Jacobi迭代方法。第三,CUDA流也被引入以在单独的流之间重叠计算。所提出的优化技术嵌入到我们的基于GPU的PCG算法。在Geforce G100上的测试结果表明,我们的SpMV内核对于超过30,0000行的大型稀疏矩阵的性能提高了近100%。同时,PCG方法的加速比也达到了7以上,使实时物理引擎成为可能。
An improved implementation of the Preconditioned Conjugate Gradient method on GPU using CUDA (Compute Unified Device Architecture) is proposed. It aims to solving the Poisson equation arising in liquid animation with high efficiency. We consider the features of the linear system obtained from the Poisson equation and propose an optimization method to solve it. First, a novel storage format called mDIA (modified diagonal storage format) is presented to improve the efficiency of the Sparse Matrix-Vector product (SpMV) operation. Second, a parallel Jacobi iterative method is proposed when using the Incomplete Cholesky preconditioner to explore inherent parallelism. Third, CUDA streams are also introduced to overlap computations among separate streams. The proposed optimization technique is embedded into our GPU based PCG algorithm. Results on Geforce G100 show that our SpMV kernel yields an improvement of nearly 100% for large sparse matrix with more than 30, 0000 rows. Also, a speedup of more than 7 is obtained for PCG method, making the real-time physics engine possible.