Accurate CUDA performance modeling for sparse matrix-vector multiplication
Accurate CUDA performance modeling for sparse matrix-vector multiplication
复制标题
DOI:
10.1109/hpcsim.2012.6266964
复制
发表时间:
2012-07
期刊:
影响因子:
--
通讯作者:
Ping Guo;Liqiang Wang
中科院分区:
文献类型:
--
作者:
Ping Guo;Liqiang Wang
This paper presents an integrated analytical and profile-based CUDA performance modeling approach to accurately predict the kernel execution times of sparse matrix-vector multiplication for CSR, ELL, COO, and HYB SpMV CUDA kernels. Based on our experiments conducted on a collection of 8 widely-used testing matrices on NVIDIA Tesla C2050, the execution times predicted by our model match the measured execution times of NVIDIA's SpMV implementations very well. Specifically, for 29 out of 32 test cases, the performance differences are under or around 7%. For the rest 3 test cases, the differences are between 8% and 10%. For CSR, ELL, COO, and HYB SpMV kernels, the differences are 4.2%, 5.2%, 1.0%, and 5.7% on the average, respectively.