A Streaming Dataflow Engine for Sparse Matrix-Vector Multiplication Using High-Level Synthesis

A Streaming Dataflow Engine for Sparse Matrix-Vector Multiplication Using High-Level Synthesis
复制标题

DOI:
10.1109/tcad.2019.2912923
复制
发表时间:
2020-06-01
影响因子:
2.9
通讯作者:
Nunez-Yanez, Jose Luis
Nunez-Yanez, Jose Luis
中科院分区:
计算机科学3区
文献类型:
--
作者:
Hosseinabady, Mohammad;Nunez-Yanez, Jose Luis

文献摘要

被引文献

相似文献

利用高级综合技术,本文提出了一种适用于嵌入式FPGA的高性能稀疏矩阵密集向量乘法(SpMV)流式并行引擎。由于SpMV是一种内存受限的算法,该引擎结合了循环流水线、流水图和数据流这三个概念,以利用FPGA可用的大部分内存带宽。本文的主要目标是证明FPGA可以为内存受限的应用程序提供与相应CPU和GPU相当的性能,但能耗明显更低。实验结果表明,FPGA提供了更高的性能相比,嵌入式GPU的中小型矩阵的平均因子为3.25,而嵌入式GPU是更快的大尺寸矩阵的平均因子为1.58。此外,与嵌入式CPU和GPU相比,FPGA实现在所考虑的矩阵范围内的能效平均为8.9倍。一个案例研究的基础上适应建议的SpMV优化,以加速支持向量机(SVM)算法,在机器学习文献中的成功分类技术之一,证明了利用建议的基于FPGA的SpMV相比,嵌入式CPU和GPU的好处。实验结果表明,FPGA的速度平均为GPU的1.7倍,消耗的能量平均为GPU的6.8倍。
Using high-level synthesis techniques, this paper proposes an adaptable high-performance streaming dataflow engine for sparse matrix dense vector multiplication (SpMV) suitable for embedded FPGAs. As the SpMV is a memory-bound algorithm, this engine combines the three concepts of loop pipelining, dataflow graph, and data streaming to utilize most of the memory bandwidth available to the FPGA. The main goal of this paper is to show that FPGAs can provide comparable performance for memory-bound applications to that of the corresponding CPUs and GPUs but with significantly less energy consumption. The experimental results indicate that the FPGA provides higher performance compared to that of embedded GPUs for small and medium-size matrices by an average factor of 3.25 whereas the embedded GPU is faster for larger size matrices by an average factor of 1.58. In addition, the FPGA implementation is more energy efficient for the range of considered matrices by an average factor of 8.9 compared to the embedded CPU and GPU. A case study based on adapting the proposed SpMV optimization to accelerate the support vector machine (SVM) algorithm, one of the successful classification techniques in the machine learning literature, justifies the benefits of utilizing the proposed FPGA-based SpMV compared to that of the embedded CPU and GPU. The experimental results show that the FPGA is faster by an average factor of 1.7 and consumes less energy by an average factor of 6.8 compared to the GPU.