FPGA-Based Acceleration for Bayesian Convolutional Neural Networks

FPGA-Based Acceleration for Bayesian Convolutional Neural Networks
复制标题

DOI:
10.1109/tcad.2022.3160948
复制
发表时间:
2022-12-01
影响因子:
2.9
通讯作者:
Luk, Wayne
Luk, Wayne
中科院分区:
计算机科学3区
文献类型:
--
作者:
Fan, Hongxiang;Ferianc, Martin;Luk, Wayne

文献摘要

被引文献

相似文献

神经网络已经在从计算机视觉(CV)到自然语言处理的各个领域展示了它们的潜力。在各种神经网络中,二维(2-D)和三维(3-D)卷积神经网络(cnn)由于其在提取二维和三维特征方面的出色能力,已被广泛应用于图像分类和视频识别等广泛的应用。然而,标准的2-D和3-D cnn无法捕捉其模型的不确定性,这对于许多安全关键应用至关重要,包括医疗保健和自动驾驶。相比之下,贝叶斯cnn (bayescnn)作为cnn的一种变体,已经证明了它们通过数学基础来表达预测中的不确定性的能力。然而,贝叶斯cnn在工业实践中并没有得到广泛的应用,因为它的计算需求来自于采样和随后在整个网络中多次前传。因此,与标准cnn相比,这些要求显著增加了计算量和内存消耗。本文提出了一种新的基于现场可编程门阵列(FPGA)的硬件架构,以加速基于蒙特卡罗dropout (MCD)的二维和三维贝叶斯cnn。与其他最先进的贝叶斯cnn加速器相比,所提出的设计可以实现高达4倍的能源效率和9倍的计算效率。提出了一个支持部分贝叶斯推理的自动框架,以探索算法和硬件性能之间的权衡。大量的实验表明,我们的框架可以有效地在设计空间中找到最优实现。
Neural networks (NNs) have demonstrated their potential in a variety of domains ranging from computer vision (CV) to natural language processing. Among various NNs, two-dimensional (2-D) and three-dimensional (3-D) convolutional NNs (CNNs) have been widely adopted for a broad spectrum of applications, such as image classification and video recognition, due to their excellent capabilities in extracting 2-D and 3-D features. However, standard 2-D and 3-D CNNs are not able to capture their model uncertainty which is crucial for many safety-critical applications, including healthcare and autonomous driving. In contrast, Bayesian CNNs (BayesCNNs), as a variant of CNNs, have demonstrated their ability to express uncertainty in their prediction via a mathematical grounding. Nevertheless, BayesCNNs have not been widely used in industrial practice due to their compute requirements stemming from sampling and subsequent forward passes through the whole network multiple times. As a result, these requirements significantly increase the amount of computation and memory consumption in comparison to standard CNNs. This article proposes a novel field-programmable gate array (FPGA)-based hardware architecture to accelerate both 2-D and 3-D BayesCNNs based on Monte Carlo dropout (MCD). Compared with other state-of-the-art accelerators for BayesCNNs, the proposed design can achieve up to four times higher energy efficiency and nine times better compute efficiency. An automatic framework capable of supporting partial Bayesian inference is proposed to explore the tradeoff between algorithm and hardware performance. Extensive experiments are conducted to demonstrate that our framework can effectively find the optimal implementations in the design space.