Improving Performance of Matrix Multiplication and FFT on GPU

Improving Performance of Matrix Multiplication and FFT on GPU
复制标题

DOI:
10.1109/icpads.2009.8
复制
发表时间:
2009-12
期刊:
2009 15th International Conference on Parallel and Distributed Systems
影响因子:
--
通讯作者:
Xiang Cui;Yifeng Chen;Hong Mei
Xiang Cui;Yifeng Chen;Hong Mei
中科院分区:
其他
文献类型:
--
作者:
Xiang Cui;Yifeng Chen;Hong Mei

文献摘要

被引文献

相似文献

本文讨论了我们在提高两个关键算法性能方面的经验:BLAS的单精度矩阵-矩阵乘法子程序(SGEMM)和CUDA的单精度FFT。前者是计算密集型的,而后者是内存带宽或通信密集型的。前者在NVIDIA GeForce GTX 280上实现了393 Gflops的峰值性能,比CUBLAS 2.0库快约5%。更好的FFT性能的结果,获得了一系列的尺寸。讨论了众核算法设计与实现的一些共同原则。
In this paper we discuss about our experiences in improving the performance of two key algorithms: the single-precision matrix-matrix multiplication subprogram (SGEMM of BLAS) and single-precision FFT using CUDA. The former is computation-intensive, while the latter is memory bandwidth or communication-intensive. A peak performance of 393 Gflops is achieved on NVIDIA GeForce GTX280 for the former, about 5% faster than the CUBLAS 2.0 library. Better FFT performance results are obtained for a range of dimensions. Some common principles are discussed for the design and implementation of many-core algorithms.