HBUNN - Hybrid Binary-Unary Neural Network: Realizing a Complete CNN on an FPGA

HBUNN - Hybrid Binary-Unary Neural Network: Realizing a Complete CNN on an FPGA
复制标题

HBUNN - 混合二元-一元神经网络:在 FPGA 上实现完整的 CNN

DOI:
10.1109/iccd46524.2019.00027
复制
发表时间:
2019
期刊:
2019 IEEE 37th International Conference on Computer Design (ICCD)
影响因子:
--
通讯作者:
K. Bazargan
K. Bazargan
中科院分区:
--
文献类型:
--
作者:
Sayed Abdolrasouol Faraji;Gaurav Singh;K. Bazargan

文献摘要

被引文献

相似文献

大多数现代神经网络以某种形式使用基本的乘法累加(MAC)运算,随着网络变得越来越大,这些大型网络的计算需求迅速增长。通常,在FPGA上实现的神经网络使用可变乘法器,以便可以在MAC中使用任何权重,但这也迫使加速器设计人员将模型的权重存储在片外存储器(DRAM)中。此外,由于当今CNN的高计算能力和高内存带宽要求,FPGA平台更难提供最佳性能。在本文中,我们提出了一种完全并行流水线的混合二进制-一进制神经网络(HBUNN)架构,以实现低成本和高性能的ResNet-18卷积神经网络。我们使用一种混合二进制-一进制方法来实现常数系数乘法器和批量归一化单元。这两个单位减少了30.7%和47.97%的硬件成本平均相比,传统的二进制等价物,分别。此外,我们提出了一种新的训练方案,使用我们的硬件成本感知正则化器,不仅提高了所提出的架构和传统的二进制架构的面积成本分别为59.3%和76.7%,但也保持相同的准确性。最后,我们使用不同的正则化器实现了三个经过训练的网络。与传统的二进制结构相比,HBUNN结构的面积成本平均降低了30%,面积×延迟成本平均降低了69%。实验结果表明,该算法的误码率为12.93%,吞吐量为278 Kfps.
Most modern neural networks use the basic Multiply-Accumulate (MAC) Operation in some form or another, and as networks get larger, the computational needs for these larger networks grow rapidly. Typically neural networks implemented on FPGAs use variable multipliers so that any weight can be used in the MAC, but this also forces the accelerator designer to store the weights of the model in off-chip memory (DRAM). Moreover, because of high computational power and high memory bandwidth requirements of today's CNNs, it is harder for FPGA platforms to deliver the best performance. In this paper, we propose a fully parallel-pipeline Hybrid Binary-Unary Neural Network (HBUNN) architecture to implement a low-cost and high-performance ResNet-18 convolutional neural network. We use a hybrid binary-unary method to implement constant-coefficient multipliers and batch normalization units. These two units reduce hardware cost by 30.7% and 47.97% on average compared to the conventional binary equivalent, respectively. Moreover, we propose a novel training scheme using our hardware cost-aware regularizers that not only improves the area cost of the proposed architecture and the conventional binary architecture by 59.3% and 76.7% respectively, but also maintains the same accuracy. Finally, we have implemented three trained networks using different regularizers. The proposed HBUNN architectures reduce the area cost by 30%, and the area × delay cost by 69% on average compared to the conventional binary architectures. The error rate of the proposed work is 12.93%, while its throughput is 278 Kfps.