SWM: A High-Performance Sparse-Winograd Matrix Multiplication CNN Accelerator

SWM: A High-Performance Sparse-Winograd Matrix Multiplication CNN Accelerator
复制标题

DOI:
10.1109/tvlsi.2021.3060041
复制
发表时间:
2021-03
影响因子:
2.8
通讯作者:
Di Wu;Xitian Fan;Wei Cao;Lingli Wang
Di Wu;Xitian Fan;Wei Cao;Lingli Wang
中科院分区:
工程技术2区
文献类型:
--
作者:
Di Wu;Xitian Fan;Wei Cao;Lingli Wang

文献摘要

相似文献

最近,许多卷积神经网络(CNN)加速器被提出,以利用网络的稀疏性来同时享受计算和存储减少的好处。然而,大多数加速器不能同时利用激活和权重的稀疏性。对于那些既利用了稀疏机会又利用了稀疏机会的工作,它们不能通过静态调度(SS)策略来实现稳定的负载均衡,该策略容易受到稀疏分布的影响。本文提出了一种均衡的压缩稀疏行格式和一种动态调度策略来改善负载均衡。提出了一种集关联结构来平衡负载均衡和硬件资源开销。我们提出SWM来加速CNN推理,它同时支持稀疏卷积和稀疏全连通(FC)层。SWM提供了对大型卷积内核的Winograd适应性,并支持16位和8位量化CNN。由于激活共享,在相同稀疏性的情况下,8位处理在理论上可以达到16位处理的两倍性能。在Xilinx VCU1525平台上实现了稀疏-Winograd卷积运算最多7.6top/S,16位量化稀疏矩阵乘法最多3top/S。对于16位量化的VGG16/ResNet50,SWM每秒可以处理310/725个图像。与最先进的作品相比,我们的设计可以实现至少1.53美元的加速比和1.8美元的能效改进。
Many convolutional neural network (CNN) accelerators are proposed to exploit the sparsity of the networks recently to enjoy the benefits of both computation and memory reduction. However, most accelerators cannot exploit the sparsity of both activations and weights. For those works that exploit both sparsity opportunities, they cannot achieve the stable load balance through a static scheduling (SS) strategy, which is vulnerable to the sparsity distribution. In this work, a balanced compressed sparse row format and a dynamic scheduling strategy are proposed to improve the load balance. A set-associate structure is also presented to tradeoff the load balance and hardware resource overhead. We propose SWM to accelerate the CNN inference, which supports both sparse convolution and sparse fully connected (FC) layers. SWM provides Winograd adaptability for large convolution kernels and supports both 16-bit and 8-bit quantized CNNs. Due to the activation sharing, 8-bit processing can achieve theoretically twice the performance of the 16-bit processing with the same sparsity. The architecture is evaluated with VGG16 and ResNet50, which achieves: at most 7.6 TOP/s for sparse-Winograd convolution and three TOP/s for sparse matrix multiplication with 16-bit quantization on Xilinx VCU1525 platform. SWM can process 310/725 images per second for VGG16/ResNet50 with 16-bit quantization. Compared with the state-of-the-art works, our design can achieve at least $1.53 \boldsymbol {\times }$ speedup and $1.8 \boldsymbol {\times }$ energy efficiency improvement.