Redesigning Triangular Dense Matrix Computations on GPUs

Redesigning Triangular Dense Matrix Computations on GPUs
复制标题

重新设计 GPU 上的三角密集矩阵计算

DOI:
10.1007/978-3-319-43659-3_35
复制
发表时间:
2016
期刊:
Parallel Comput.
影响因子:
--
通讯作者:
D. Keyes
D. Keyes
中科院分区:
--
文献类型:
--
作者:
A. Charara;H. Ltaief;D. Keyes

文献摘要

被引文献

相似文献

在GPU硬件加速器上实现了一种新的三角矩阵-矩阵乘法TRMM和三角求解TRSM核。虽然是Level 3 BLAS家族的一部分,但这些高度计算密集型内核无法在GPU上实现理论峰值性能的百分比,这是在硬件加速器上运行具有类似表面积与体积比的内核时所期望的,即,标准的矩阵-矩阵乘法GEMM。作者建议采用递归公式,这丰富了TRMM和TRSM内部结构与GEMM调用,因此,减少内存流量,同时提高并发水平。新的实现能够有效地使用GPU内存层次结构,并减轻延迟开销,以更高缓存级别的速度运行。性能比较显示,在不同GPU代次中,与NVIDIA cuBLAS现有的最先进TRMM和TRSM实现相比,大型密集矩阵尺寸的加速分别提高了3/4至8倍和2倍。一旦集成到基于高层次Cholesky的稠密线性代数算法中,对整体应用程序的性能影响分别与与cuBLAS TRMM和TRSM内核链接的等效本地实现相比,表现出高达四倍和两倍的加速比。新的TRMM/TRSM内核实现是开源KBLAS软件库http://ecrc.kaust.edu.sa/Pages/Res-kblas.aspx的一部分,并将在即将发布的v8.0版本中集成到NVIDIA cuBLAS库中。
A new implementation of the triangular matrix-matrix multiplication TRMM and the triangular solve TRSM kernels are described on GPU hardware accelerators. Although part of the Level 3 BLAS family, these highly computationally intensive kernels fail to achieve the percentage of the theoretical peak performance on GPUs that one would expect when running kernels with similar surface-to-volume ratio on hardware accelerators, i.e., the standard matrix-matrix multiplication GEMM. The authors propose adopting a recursive formulation, which enriches the TRMM and TRSM inner structures with GEMM calls and, therefore, reduces memory traffic while increasing the level of concurrency. The new implementation enables efficient use of the GPU memory hierarchy and mitigates the latency overhead, to run at the speed of the higher cache levels. Performance comparisons show upi¾źto eightfold and twofold speedups for large dense matrix sizes, against the existing state-of-the-art TRMM and TRSM implementations from NVIDIA cuBLAS, respectively, across various GPU generations. Once integrated into high-level Cholesky-based dense linear algebra algorithms, the performance impact on the overall applications demonstrates upi¾źto fourfold and twofold speedups, against the equivalent native implementations, linked with cuBLAS TRMM and TRSM kernels, respectively. The new TRMM/TRSM kernel implementations are part of the open-source KBLAS software library http://ecrc.kaust.edu.sa/Pages/Res-kblas.aspx and are lined up for integration into the NVIDIA cuBLAS library in the upcoming v8.0 release.