Fast-BCNN: Massive Neuron Skipping in Bayesian Convolutional Neural Networks

Fast-BCNN: Massive Neuron Skipping in Bayesian Convolutional Neural Networks
复制标题

DOI:
10.1109/micro50266.2020.00030
复制
发表时间:
2020-10
期刊:
2020 53rd Annual IEEE/ACM International Symposium on Microarchitecture (MICRO)
影响因子:
--
通讯作者:
Qiyu Wan;Xin Fu
Qiyu Wan;Xin Fu
中科院分区:
其他
文献类型:
--
作者:
Qiyu Wan;Xin Fu

文献摘要

被引文献

相似文献

贝叶斯卷积神经网络(Bayesian Convolutional Neural Networks,BCNN)是卷积神经网络的一种鲁棒形式,具有不确定性估计能力。通过在原始CNN中的每个卷积层之后添加dropout层来实现BCNN模型。通过多次执行随机推理,BCNN能够提供反映最终预测不确定性的输出分布。在这个过程中重复的推理导致更长的执行时间,这使得在现实世界的应用中将贝叶斯技术应用于CNN具有挑战性。在这项研究中,我们提出了Fast-BCNN,这是一种基于FPGA的硬件加速器设计,可以在重复的BCNN推理过程中智能地跳过两种类型神经元的冗余计算。首先,在一个样本推理中,我们的目标是跳过由dropout掩码预先确定的丢弃的神经元。其次,通过利用来自第一个推理和dropout掩码的信息,我们预测零神经元,并在接下来的样本推理中跳过所有相应的计算。特别地,采用优化算法来保证零神经元预测的准确性,同时实现最大的计算减少。为了在硬件层面支持我们的神经元跳过策略,我们探索了CNN卷积的有效并行性,以优雅地跳过两种类型神经元的相应计算,然后我们提出了一种新的PE架构,该架构可以以可忽略的开销来适应卷积和预测的并行操作。实验结果表明,我们的Fast-BCNN实现了2.1~8.2倍的加速比和44%~84%的能量减少比基线CNN加速器。
Bayesian Convolutional Neural Networks (BCNNs) have emerged as a robust form of Convolutional Neural Networks (CNNs) with the capability of uncertainty estimation. A BCNN model is implemented by adding a dropout layer after each convolutional layer in the original CNN. By executing the stochastic inferences many times, BCNNs are able to provide an output distribution that reflects the uncertainty of the final prediction. Repeated inferences in this process lead to much longer execution time, which makes it challenging to apply Bayesian technique to CNNs in real-world applications. In this study, we propose Fast-BCNN, an FPGA-based hardware accelerator design that intelligently skips the redundant computations for two types of neurons during repeated BCNN inferences. Firstly, within a sample inference, we aim to skip the dropped neurons that predetermined by dropout masks. Secondly, by leveraging the information from the first inference and dropout masks, we predict the zero neurons and skip all their corresponding computations during the following sample inferences. Particularly, an optimization algorithm is employed to guarantee the accuracy of zero neuron prediction while achieving the maximal computation reduction. To support our neuron skipping strategy at hardware level, we explore an efficient parallelism for CNN convolution to gracefully skip the corresponding computations for both types of neurons, we then propose a novel PE architecture that accommodates the parallel operation of convolution and prediction with negligible overhead. Experimental results demonstrate that our Fast-BCNN achieves 2.1~8.2× speedup and 44%~84% energy reduction over the baseline CNN accelerator.