Residue-Net: Multiplication-free Neural Network by In-situ No-loss Migration to Residue Number Systems

Residue-Net: Multiplication-free Neural Network by In-situ No-loss Migration to Residue Number Systems
复制标题

DOI:
10.1145/3394885.3431541
复制
发表时间:
2021-01
期刊:
2021 26th Asia and South Pacific Design Automation Conference (ASP-DAC)
影响因子:
--
通讯作者:
Sahand Salamat;Sumiran Shubhi;Behnam Khaleghi;Tajana Simunic
Sahand Salamat;Sumiran Shubhi;Behnam Khaleghi;Tajana Simunic
中科院分区:
其他
文献类型:
--
作者:
Sahand Salamat;Sumiran Shubhi;Behnam Khaleghi;Tajana Simunic

文献摘要

被引文献

相似文献

深度神经网络被广泛地部署在嵌入式设备上,以解决从边缘感知到自动驾驶的广泛问题。这些网络的精确度通常与其复杂性成正比。量化模型参数(即,权重)和/或激活以在保持精度的同时减轻这些网络的复杂性是一种流行的强大技术。尽管如此,以前的研究表明,量化水平是有限的,因为网络的精度随后会下降。我们提出了一种无乘法的神经网络加速器--残数网络,它使用残数系统(RNS)来实现大幅度的能量减少。RNS将操作分解为几个更容易实现的较小操作。此外,留数网用非复杂的、节能的移位和加法运算取代了大量昂贵的乘法运算,进一步简化了神经网络的计算复杂性。为了评估我们提出的加速器的效率,我们将残差网络的性能与四个广泛使用的网络,即LeNet、AlexNet、VGG16和ResNet-50的基线FPGA实现进行了比较。当提供与基准相同的性能时,残差网络平均将面积和功率(即能量)分别减少36%和23%,而不会造成精度损失。利用节省的面积对量化后的RNS网络进行并行加速,使其吞吐量提高2.8倍,能量提高2.7倍。
Deep eural networks are widely deployed on embedded devices to solve a wide range of problems from edge-sensing to autonomous driving. The accuracy of these networks is usually proportional to their complexity. Quantization of model parameters (i.e., weights) and/or activations to alleviate the complexity of these networks while preserving accuracy is a popular powerful technique. Nonetheless, previous studies have shown that quantization level is limited as the accuracy of the network decreases afterward. We propose Residue-Net, a multiplication-free accelerator for neural networks that uses Residue Number System (RNS) to achieve substantial energy reduction. RNS breaks down the operations to several smaller operations that are simpler to implement. Moreover, Residue-Net replaces the copious of costly multiplications with non-complex, energy-efficient shift and add operations to further simplify the computational complexity of neural networks. To evaluate the efficiency of our proposed accelerator, we compared the performance of Residue-Net with a baseline FPGA implementation of four widely-used networks, viz., LeNet, AlexNet, VGG16, and ResNet-50. When delivering the same performance as the baseline, Residue-Net reduces the area and power (hence energy) respectively by 36% and 23%, on average with no accuracy loss. Leveraging the saved area to accelerate the quantized RNS network through parallelism, Residue-Net improves its throughput by 2.8× and energy by 2.7×.