A deeply-pipelined FPGA-based SpMV accelerator with a hardware-friendly storage scheme

A deeply-pipelined FPGA-based SpMV accelerator with a hardware-friendly storage scheme
复制标题

DOI:
10.1587/elex.12.20150161
复制
发表时间:
2015-05
期刊:
IEICE Electron. Express
影响因子:
--
通讯作者:
Song Guo;Y. Dou;Yuanwu Lei;Guiming Wu
Song Guo;Y. Dou;Yuanwu Lei;Guiming Wu
中科院分区:
其他
文献类型:
--
作者:
Song Guo;Y. Dou;Yuanwu Lei;Guiming Wu

文献摘要

被引文献

相似文献

提出了一种基于现场可编程门阵列(FPGA)的高性能稀疏矩阵向量乘法(SpMV)加速器。通过采用一种硬件友好的存储方案--可变位宽坐标块准压缩稀疏行,通过嵌套块压缩和可变位宽列索引编码方案,可以大大减少冗余计算和内存访问。在Xilinx公司的Virtex XC 7VX 485 T FPGA平台上实现了一个深度流水线的SpMV加速器,该加速器可以处理任意大小和稀疏模式的稀疏矩阵。实验结果表明,与Convey平台(HC-1和HC-2 ex)和Nvidia Tesla S1070 GPU平台相比,该设计在大多数测试矩阵上都能获得更高的性能,内存带宽利用率提高了13倍。
This paper presents a high performance sparse matrix-vector multiplication (SpMV) accelerator on the field-programming gate array (FPGA). By exploiting a hardware-friendly storage scheme, named as Variable-Bit-Width Coordinate Block Quasi Compressed Sparse Row, the redundant computation and memory accesses can be reduced greatly through the nested block compression and variable-bit-width column-index encoding schemes. Based on the proposed compression scheme, a deeply-pipelined SpMV accelerator is implemented on a Xilinx Virtex XC7VX485T FPGA platform, which can handle sparse matrices with arbitrary size and sparsity pattern. Experimental results show that the proposed design can gain higher performance for most of the tested matrices and improve the utilization of the memory bandwidth up to 13×, compared with the previous works on the Convey platforms (HC-1 and HC-2ex) and Nvidia Tesla S1070 GPU platform.