A Cross-Platform SpMV Framework on Many-Core Architectures

A Cross-Platform SpMV Framework on Many-Core Architectures
复制标题

DOI:
10.1145/2994148
复制
发表时间:
2016-10
期刊:
ACM Transactions on Architecture and Code Optimization (TACO)
影响因子:
--
通讯作者:
Yunquan Zhang;Shigang Li;Shengen Yan;Huiyang Zhou
Yunquan Zhang;Shigang Li;Shengen Yan;Huiyang Zhou
中科院分区:
其他
文献类型:
--
作者:
Yunquan Zhang;Shigang Li;Shengen Yan;Huiyang Zhou

文献摘要

被引文献

相似文献

稀疏矩阵向量乘法(Sparse Matrix-Vector Multiplication,SpMV)是工程计算和科学计算中的一种重要运算。虽然以前的工作已经在众核架构上优化SpMV方面取得了令人印象深刻的进展,但负载不平衡和高内存带宽仍然是关键的性能瓶颈。我们提出了我们的新的解决方案,这些问题,为GPU和英特尔MIC众核架构。首先,我们设计了一种新的SpMV格式,称为分块压缩公共坐标(BCCOO)。BCCOO通过使用位标志来存储行索引来扩展阻塞的公共坐标(COO),以缓解带宽问题。我们通过将矩阵划分为垂直切片以获得更好的数据局部性来进一步改进这种格式。然后,为了解决负载不平衡的问题,我们提出了一个高效的基于矩阵的分段求和/扫描算法的SpMV,消除了全局同步。最后,我们介绍了一个自动调整框架来选择优化参数。实验结果表明,我们提出的框架具有显着的优势,现有的SpMV库。在单精度下,我们提出的方案在AMD FirePro W8000上的性能平均优于clSpMV COCKTAIL格式255%,在GeForce Titan X上的性能平均优于CSPARSE V7.0 73.7%,优于CSR 5 53.6%;在双精度下,我们提出的方案在Tesla K20上的性能平均优于RISTOPARSE V7.0 34.0%,平均优于CSR 5 16.2%,与Intel MIC上的CSR 5性能相当。
Sparse Matrix-Vector multiplication (SpMV) is a key operation in engineering and scientific computing. Although the previous work has shown impressive progress in optimizing SpMV on many-core architectures, load imbalance and high memory bandwidth remain the critical performance bottlenecks. We present our novel solutions to these problems, for both GPUs and Intel MIC many-core architectures. First, we devise a new SpMV format, called Blocked Compressed Common Coordinate (BCCOO). BCCOO extends the blocked Common Coordinate (COO) by using bit flags to store the row indices to alleviate the bandwidth problem. We further improve this format by partitioning the matrix into vertical slices for better data locality. Then, to address the load imbalance problem, we propose a highly efficient matrix-based segmented sum/scan algorithm for SpMV, which eliminates global synchronization. At last, we introduce an autotuning framework to choose optimization parameters. Experimental results show that our proposed framework has a significant advantage over the existing SpMV libraries. In single precision, our proposed scheme outperforms clSpMV COCKTAIL format by 255% on average on AMD FirePro W8000, and outperforms CUSPARSE V7.0 by 73.7% on average and outperforms CSR5 by 53.6% on average on GeForce Titan X; in double precision, our proposed scheme outperforms CUSPARSE V7.0 by 34.0% on average and outperforms CSR5 by 16.2% on average on Tesla K20, and has equivalent performance compared with CSR5 on Intel MIC.