Bisection Neural Network Toward Reconfigurable Hardware Implementation

Bisection Neural Network Toward Reconfigurable Hardware Implementation
复制标题

DOI:
10.1109/tnnls.2022.3195821
复制
发表时间:
2022-08
影响因子:
10.4
通讯作者:
Yan Chen;Renyuan Zhang;Yirong Kan;Sa Yang;Y. Nakashima
Yan Chen;Renyuan Zhang;Yirong Kan;Sa Yang;Y. Nakashima
中科院分区:
计算机科学1区
文献类型:
--
作者:
Yan Chen;Renyuan Zhang;Yirong Kan;Sa Yang;Y. Nakashima

文献摘要

相似文献

本文提出了一种硬件友好的二分神经网络(BNN)拓扑结构,用于在任意片上配置下近似实现大量复杂功能。与传统的可重构全连接神经网络(FC-NN)电路拓扑结构不同,硬件友好的拓扑结构执行NN行为,其中每个神经元包括输入和输出的两个恒定突触连接。与FC-NN相比,BNN电路拓扑的重新配置消除了硬件中大量的虚拟突触连接。作为主要的目标应用,这项工作旨在建立一个通用的BNN电路拓扑,提供大量的NN回归。为了实现这一目标,我们证明了FC-NN电路拓扑的神经网络行为可以等价地迁移到BNN电路拓扑。我们引入了两种方法,包括精化训练算法和倒金字塔策略,以进一步减少神经元和突触的数量。最后,我们进行了误差容差分析,为超高效的硬件实现提供了指导。与目前最先进的基于FC-NN电路拓扑的TrueNorth基线相比,该设计可以实现17.8-22.2倍的硬件开销和小于1%的误差。
A hardware-friendly bisection neural network (BNN) topology is proposed in this work for approximately implementing massive pieces of complex functions in arbitrary on-chip configurations. Instead of the conventional reconfigurable fully connected neural network (FC-NN) circuit topology, the proposed hardware-friendly topology performs NN behaviors in a bisection structure, in which each neuron includes two constant synapse connections for both inputs and outputs. Compared with the FC-NN one, the reconfiguration of the BNN circuit topology eliminates the remarkable amount of dummy synapse connections in hardware. As the main target application, this work aims at building a general-purpose BNN circuit topology that offers a great amount of NN regressions. To achieve this target, we prove that the NN behaviors of the FC-NN circuit topologies can be migrated to the BNN circuit topologies equivalently. We introduce two approaches including the refining training algorithm and the inverted-pyramidal strategy to further reduce the number of neurons and synapses. Finally, we conduct the inaccuracy tolerance analysis to suggest the guideline for ultra-efficient hardware implementations. Compared with the state-of-the-art FC-NN circuit topology-based TrueNorth baseline, the proposed design can achieve 17.8– $22.2\times $ hardware reduction and less than 1% inaccuracy.