Sparse and dense matrix multiplication hardware for heterogeneous multi-precision neural networks

Sparse and dense matrix multiplication hardware for heterogeneous multi-precision neural networks
复制标题

DOI:
10.1016/j.array.2021.100101
复制
发表时间:
2021-11
期刊:
影响因子:
--
通讯作者:
J. Núñez-Yáñez;Mohammad Hosseinabady
J. Núñez-Yáñez;Mohammad Hosseinabady
中科院分区:
--
文献类型:
--
作者:
J. Núñez-Yáñez;Mohammad Hosseinabady

文献摘要

相似文献

在本文中,我们提出了硬件加速器创建高层次的合成技术,稀疏和密集矩阵乘法运算。这些内核可以以不同的精度运行,并被设计为集成在异构CPU-FPGA系统中,用于边缘AI应用。该方法涉及量化稀疏意识的训练,它被应用到一个案例研究,包括人类活动分类。我们首先研究了量化和稀疏性对神经网络准确性的影响,当存在递归层时,卷积层、密集层和递归层对修剪有更好的耐受性。然后,我们提出了硬件加速器,可以在运行时切换精度,并与任何矩阵大小的最大配置在编译时。我们比较了这些加速器在不同精度和稀疏度级别下的性能,并创建了一个性能模型来实现工作负载平衡。结果表明,当稀疏度大于70%时,稀疏矩阵乘子的性能优于稠密矩阵乘子,且当采用更高精度的算法或结构剪枝时,这种改进更为明显.此外,高达99%的稀疏度水平可以保持网络所需的准确度,特别是在部署循环层时。总体而言,稀疏和密集性能之间的平衡取决于矩阵形状、精度、结构修剪和稀疏级别,并且性能建模可以用于平衡异构配置中的并发执行。
In this paper, we present hardware accelerators created with high-level synthesis techniques for sparse and dense matrix multiplication operations. The cores can operate with different precisions and are designed to be integrated in a heterogeneous CPU-FPGA system for Edge AI applications. The methodology involves quantization-sparsity aware training and it is applied to a case study consisting of human activity classification. We initially investigate the effects of quantization and sparsity on the accuracy of neural networks with convolution, dense and recurrent layers observing better tolerance to pruning when recurrent layers are present. Then, we propose the hardware accelerators that can switch precision at run-time and work with any matrix size up to a maximum configured at compile time. We compare the performance of these accelerators at different levels of precision and sparsity levels and create a performance model to enable workload balancing. The results show that the proposed sparse matrix multipliers can outperform dense multipliers when sparsity levels are higher than 70% and the improvements are more evident when higher precision arithmetic or structural pruning is used. Additionally, sparsity levels as high as 99% can maintain the level of accuracy required in the network especially when recurrent layers are deployed. Overall, the balance between sparse and dense performance depends on matrix shape, precision, structural pruning and sparsity levels and performance modelling can be used to balance concurrent execution in a heterogeneous configuration.