Redesigning Triangular Dense Matrix Computations on GPUs
Redesigning Triangular Dense Matrix Computations on GPUs
复制标题
重新设计 GPU 上的三角密集矩阵计算
DOI:
10.1007/978-3-319-43659-3_35
复制
发表时间:
2016
期刊:
影响因子:
--
通讯作者:
D. Keyes
中科院分区:
文献类型:
--
作者:
A. Charara;H. Ltaief;D. Keyes
A new implementation of the triangular matrix-matrix multiplication TRMM and the triangular solve TRSM kernels are described on GPU hardware accelerators. Although part of the Level 3 BLAS family, these highly computationally intensive kernels fail to achieve the percentage of the theoretical peak performance on GPUs that one would expect when running kernels with similar surface-to-volume ratio on hardware accelerators, i.e., the standard matrix-matrix multiplication GEMM. The authors propose adopting a recursive formulation, which enriches the TRMM and TRSM inner structures with GEMM calls and, therefore, reduces memory traffic while increasing the level of concurrency. The new implementation enables efficient use of the GPU memory hierarchy and mitigates the latency overhead, to run at the speed of the higher cache levels. Performance comparisons show upi¾źto eightfold and twofold speedups for large dense matrix sizes, against the existing state-of-the-art TRMM and TRSM implementations from NVIDIA cuBLAS, respectively, across various GPU generations. Once integrated into high-level Cholesky-based dense linear algebra algorithms, the performance impact on the overall applications demonstrates upi¾źto fourfold and twofold speedups, against the equivalent native implementations, linked with cuBLAS TRMM and TRSM kernels, respectively. The new TRMM/TRSM kernel implementations are part of the open-source KBLAS software library http://ecrc.kaust.edu.sa/Pages/Res-kblas.aspx and are lined up for integration into the NVIDIA cuBLAS library in the upcoming v8.0 release.