O3BNN-R: An Out-of-Order Architecture for High-Performance and Regularized BNN Inference

O3BNN-R: An Out-of-Order Architecture for High-Performance and Regularized BNN Inference
复制标题

DOI:
10.1109/tpds.2020.3013637
复制
发表时间:
2021-01
影响因子:
5.3
通讯作者:
Tong Geng;Ang Li;Tianqi Wang;Chunshu Wu;Yanfei Li;Runbin Shi;Wei Wu;M. Herbordt
Tong Geng;Ang Li;Tianqi Wang;Chunshu Wu;Yanfei Li;Runbin Shi;Wei Wu;M. Herbordt
中科院分区:
计算机科学2区
文献类型:
--
作者:
Tong Geng;Ang Li;Tianqi Wang;Chunshu Wu;Yanfei Li;Runbin Shi;Wei Wu;M. Herbordt

文献摘要

被引文献

相似文献

二进制的神经网络(BNN)大大降低了计算复杂性和内存需求,在成本和功率限制域(例如IoT和Smart Edge-evices)中显示出潜力在本文中,我们证明了高度符号的BNN模型可以通过基于BNN特异性特征(OOO)体系结构的两个新观测来动态修剪不规则的冗余边缘。在推断期间可以在运行时确定神经元的二进制输出的情况下,可以降低边缘评估。为了进一步增强修剪机会,我们进行了一种算法/体系结构共同设计方法,在训练阶段,我们使用包括嵌入式FPGA的FPGA进行了使用,以使用vgg-16的网络来增强训练阶段的损失功能。 ,对于Imagenet的Alexnet和CIFAR-10的VGG样网络表明,没有调节的O3BNN-R可以平均修剪30%的操作,而无需任何精确损失,对FPGA/GPU/CPU的最先进的BNN实施,平均有34倍的能源效率。
Binarized Neural Networks (BNN), which significantly reduce computational complexity and memory demand, have shown potential in cost- and power-restricted domains, such as IoT and smart edge-devices, where reaching certain accuracy bars is sufficient and real-time is highly desired. In this article, we demonstrate that the highly-condensed BNN model can be shrunk significantly by dynamically pruning irregular redundant edges. Based on two new observations on BNN-specific properties, an out-of-order (OoO) architecture, O3BNN-R, which can curtail edge evaluation in cases where the binary output of a neuron can be determined early at runtime during inference, is proposed. Similar to instruction level parallelism (ILP), fine-grained, irregular, and runtime pruning opportunities are traditionally presumed to be difficult to exploit. To further enhance the pruning opportunities, we conduct an algorithm/architecture co-design approach where we augment the loss function during the training stage with specialized regularization terms favoring edge pruning. We evaluate our design on an embedded FPGA using networks that include VGG-16, AlexNet for ImageNet, and a VGG-like network for Cifar-10. Results show that O3BNN-R without regularization can prune, on average, 30 percent of the operations, without any accuracy loss, bringing 2.2× inference-speedup, and on average 34× energy-efficiency improvement over state-of-the-art BNN implementations on FPGA/GPU/CPU. With regularization at training, the performance is further improved, on average, by 15 percent.