Basic Linear Algebra Operations on TensorCore GPU

Basic Linear Algebra Operations on TensorCore GPU
复制标题

TensorCore GPU 上的基本线性代数运算

DOI:
--
复制
发表时间:
2020
期刊:
ACM SIGPLAN Symposium on Scala
影响因子:
--
通讯作者:
Panruo Wu
Panruo Wu
中科院分区:
--
文献类型:
--
作者:
Shaoshuai Zhang;Vivek Karihaloo;Panruo Wu

文献摘要

被引文献

相似文献

受高速矩阵计算和训练深度神经网络的需求的鼓舞,NVIDIA图形处理器中引入了TensorCore来进一步加速矩阵-矩阵乘法。它支持非常快的半精度通用矩阵乘法(GEM),比单精度CUDA核心GEM快约8倍。到目前为止,使用TensorCore图形处理器进行矩阵运算,而不是矩阵-矩阵乘法,还处于开发阶段。在本文中,我们提出了一些有效的利用TensorCore的BLAS3操作。实验结果表明,该算法的性能优于CUBLAS对应的例程和朴素的TensorCore实现,加速比高达4.7倍。
Encouraged by the requirement of high speed matrix computations and training deep neural networks, TensorCore was introduced in NVIDIA GPU to further accelerate matrix-matrix multiplication. It supports very fast half precision general matrix matrix multiplications (GEMMs), which is around 8x faster than single precision CUDA core GEMMs. So far the use of TensorCore GPU for matrix operations other than matrix-matrix multiplications is under developed. In this paper, we propose some efficient BLAS3 operations that exploits TensorCore. The experimental results show that the proposed algorithms outperform cublas corresponding routines and the naive TensorCore implementation with up to 4.7x speedup.