Accelerating Bayesian Neural Networks via Algorithmic and Hardware Optimizations

Accelerating Bayesian Neural Networks via Algorithmic and Hardware Optimizations
复制标题

DOI:
10.1109/tpds.2022.3153682
复制
发表时间:
2022
影响因子:
5.3
通讯作者:
Hongxiang Fan;Martin Ferianc;Zhiqiang Que;Xinyu Niu;Miguel L. Rodrigues;Wayne Luk
Hongxiang Fan;Martin Ferianc;Zhiqiang Que;Xinyu Niu;Miguel L. Rodrigues;Wayne Luk
中科院分区:
计算机科学2区
文献类型:
--
作者:
Hongxiang Fan;Martin Ferianc;Zhiqiang Que;Xinyu Niu;Miguel L. Rodrigues;Wayne Luk

文献摘要

被引文献

相似文献

贝叶斯神经网络(BayesNN)已经在各种安全关键应用中展示了其优势,例如自动驾驶或医疗保健,因为它们能够捕获和表示模型不确定性。然而,标准的贝叶斯神经网络需要重复运行,因为蒙特卡洛采样来量化它们的不确定性,这给它们的实际硬件性能带来了负担。为了解决这个性能问题,本文系统地利用广泛的结构稀疏性和冗余计算贝叶斯神经网络。与标准卷积神经网络中存在的非结构化或结构化稀疏性不同,贝叶斯神经网络的结构化稀疏性是通过Monte Carlo Dropout及其在不确定性估计和预测期间所需的相关采样来引入的,可以通过算法和硬件优化来利用。我们首先将观察到的稀疏模式分为三类:辍学稀疏,层稀疏和样本稀疏。在算法方面,提出了一个框架,自动探索这三个稀疏类别,而不牺牲算法性能。我们证明了结构化稀疏性可以将CPU设计加速高达49倍,GPU设计加速高达40倍。在硬件方面,提出了一种新的硬件架构来加速贝叶斯神经网络,该架构使用运行时自适应硬件引擎和智能跳过支持来实现高硬件性能。在FPGA上实现所提出的硬件设计,我们的实验表明,算法优化的贝叶斯神经网络可以实现高达56倍的加速比相比,未优化的贝叶斯网络。与优化的GPU实现相比,我们的FPGA设计实现了高达7.6倍的加速比和高达39.3倍的能效
Bayesian neural networks (BayesNNs) have demonstrated their advantages in various safety-critical applications, such as autonomous driving or healthcare, due to their ability to capture and represent model uncertainty. However, standard BayesNNs require to be repeatedly run because of Monte Carlo sampling to quantify their uncertainty, which puts a burden on their real-world hardware performance. To address this performance issue, this paper systematically exploits the extensive structured sparsity and redundant computation in BayesNNs. Different from the unstructured or structured sparsity existing in standard convolutional NNs, the structured sparsity of BayesNNs is introduced by Monte Carlo Dropout and its associated sampling required during uncertainty estimation and prediction, which can be exploited through both algorithmic and hardware optimizations. We first classify the observed sparsity patterns into three categories: dropout sparsity, layer sparsity and sample sparsity. On the algorithmic side, a framework is proposed to automatically explore these three sparsity categories without sacrificing algorithmic performance. We demonstrated that structured sparsity can be exploited to accelerate CPU designs by up to 49 times, and GPU designs by up to 40 times. On the hardware side, a novel hardware architecture is proposed to accelerate BayesNNs, which achieves a high hardware performance using the runtime adaptable hardware engines and the intelligent skipping support. Upon implementing the proposed hardware design on an FPGA, our experiments demonstrated that the algorithm-optimized BayesNNs can achieve up to 56 times speedup when compared with unoptimized Bayesian nets. Comparing with the optimized GPU implementation, our FPGA design achieved up to 7.6 times speedup and up to 39.3 times higher energy efficiency