Basic Linear Algebra Operations on TensorCore GPU
Basic Linear Algebra Operations on TensorCore GPU
复制标题
TensorCore GPU 上的基本线性代数运算
DOI:
--
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Panruo Wu
中科院分区:
文献类型:
--
作者:
Shaoshuai Zhang;Vivek Karihaloo;Panruo Wu
Encouraged by the requirement of high speed matrix computations and training deep neural networks, TensorCore was introduced in NVIDIA GPU to further accelerate matrix-matrix multiplication. It supports very fast half precision general matrix matrix multiplications (GEMMs), which is around 8x faster than single precision CUDA core GEMMs. So far the use of TensorCore GPU for matrix operations other than matrix-matrix multiplications is under developed. In this paper, we propose some efficient BLAS3 operations that exploits TensorCore. The experimental results show that the proposed algorithms outperform cublas corresponding routines and the naive TensorCore implementation with up to 4.7x speedup.