Fast Implementation of General Matrix-Vector Multiplication (GEMV) on Kepler GPUs

Fast Implementation of General Matrix-Vector Multiplication (GEMV) on Kepler GPUs
复制标题

DOI:
10.1109/pdp.2015.66
复制
发表时间:
2015-03
期刊:
2015 23rd Euromicro International Conference on Parallel, Distributed, and Network-Based Processing
影响因子:
--
通讯作者:
Daichi Mukunoki;Toshiyuki Imamura;D. Takahashi
Daichi Mukunoki;Toshiyuki Imamura;D. Takahashi
中科院分区:
其他
文献类型:
--
作者:
Daichi Mukunoki;Toshiyuki Imamura;D. Takahashi

文献摘要

相似文献

针对NVIDIA开普勒体系结构图形处理器(GPU)上的列主和非转置矩阵,提出了一种通用矩阵向量乘(GEMV)例程的快速实现方法。我们首先使用共享内存和寄存器的典型阻塞技术以及128位向量加载/存储指令来实现GEMV内核。在我们的初步调查中,我们发现,即使内核在某些矩阵大小下可以接近实际的峰值GPU吞吐量,性能也会根据问题的大小而周期性地波动。在我们的下一步中,我们使用基于线程块调度机制的性能模型来研究波动的原因,然后创建了一种确定最优线程块大小的方法来避免这些波动。结果表明,当运行在两个开普勒架构的GPU上时,我们的单精度GEMV(SGEMV)例程在吞吐量和性能稳定性(相对于问题大小)方面都取得了比现有实现:CUBLAS 6.5、MAGMA 1.4.1和KBLAS 1.0更好的性能。我们的实现技术不仅可以用于SGEMV,还可以用于双精度(DGEMV)、单复数(CGEMV)和双复数(ZGEMV)。在主要讨论开普勒体系结构的同时,我们还探讨了方案在下一代开普勒体系结构Maxwell体系结构上的实现性能。
This paper proposes a fast implementation method for the general matrix-vector multiplication (GEMV) routine, which is one of the level-2 Basic Linear Algebra Subprograms (BLAS) subroutines, for a column-major and non-transposed matrix on NVIDIA Kepler architecture graphics processing units (GPUs). We began by implementing the GEMV kernel using typical blocking techniques for shared-memory and register along with 128-bit vector load/store instructions. In our initial investigation, we found that even though the kernel could approach actual peak GPU throughput at some matrix sizes, performance fluctuates periodically depending on the problem size. In our next step, we investigated the reason for the fluctuations using a performance model based on a thread-block scheduling mechanism, and then created a method of determining optimal thread-block sizes that avoids those fluctuations. As the results show, when run on two Kepler architecture GPUs, our single-precision GEMV (SGEMV) routine achieved better performance in terms of both throughput and performance stability (with respect to the problem size) when compared to existing implementations: CUBLAS 6.5, MAGMA 1.4.1 and KBLAS 1.0. Our implementation techniques can be used not only for SGEMV but also double-precision (DGEMV), single-complex (CGEMV), and double-complex (ZGEMV). While this paper discusses primarily Kepler architecture, we also explore the performance of proposal implementation on Maxwell architecture, which is the next generation of Kepler architecture.