High-Performance Matrix-Vector Multiplication on the GPU

High-Performance Matrix-Vector Multiplication on the GPU
复制标题

GPU 上的高性能矩阵向量乘法

DOI:
--
复制
发表时间:
2011
期刊:
Euro-Par Workshops
影响因子:
--
通讯作者:
H. H. Sørensen
H. H. Sørensen
中科院分区:
--
文献类型:
--
作者:
H. H. Sørensen

文献摘要

被引文献

相似文献

在本文中,我们开发了一个高性能的GPU内核的最流行的密集线性代数运算之一,矩阵向量乘法。目标硬件是最新的Nvidia Tesla 20系列(Fermi架构),它是为科学计算而设计的。我们表明,它本质上是一个问题,充分利用细粒度的并行性的众核GPU,以实现高性能的密集矩阵向量乘法。我们表明,自动调整可以成功地采用GPU内核,使它表现良好的所有矩阵的形状和大小。
In this paper, we develop a high-performance GPU kernel for one of the most popular dense linear algebra operations, the matrix-vector multiplication. The target hardware is the most recent Nvidia Tesla 20-series (Fermi architecture), which is designed from the ground up for scientific computing. We show that it is essentially a matter of fully utilizing the fine-grained parallelism of the many-core GPU in order to achieve high-performance for dense matrix-vector multiplication. We show that auto-tuning can be successfully employed to the GPU kernel so that it performs well for all matrix shapes and sizes.