Uncertainty propagation for dropout-based Bayesian neural networks

Uncertainty propagation for dropout-based Bayesian neural networks
复制标题

DOI:
10.1016/j.neunet.2021.09.005
复制
发表时间:
2021-09
期刊:
Neural networks : the official journal of the International Neural Network Society
影响因子:
--
通讯作者:
Yuki Mae;Wataru Kumagai;T. Kanamori
Yuki Mae;Wataru Kumagai;T. Kanamori
中科院分区:
其他
文献类型:
--
作者:
Yuki Mae;Wataru Kumagai;T. Kanamori

文献摘要

相似文献

当深度神经网络(DNN)应用于实际问题时,不确定性评估是一项核心技术。在实际应用中,我们经常会遇到训练过程中未曾见过的意外样本。不仅要达到较高的预测精度,而且要检测出不确定的数据,这对安全关键系统具有重要意义。在统计学和机器学习中,贝叶斯推理被用于不确定性评估。贝叶斯神经网络(BNN)最近在这方面引起了相当大的关注,因为使用丢弃训练的DNN被解释为贝叶斯方法。在此基础上,提出了计算离散神经网络贝叶斯预测分布的几种方法。虽然蒙特卡罗方法MC Dropout是一种流行的不确定度评估方法,但它需要对具有随机抽样的权重参数的DNN进行多次重复前馈计算。为了克服计算上的问题,我们提出了一种无采样的方法来评估不确定性。我们的方法将使用丢弃训练的神经网络转换为相应的具有方差传播的贝叶斯神经网络。该方法不仅适用于前馈神经网络,也适用于LSTM等递归神经网络。在用RNN进行语言建模的数值实验中,我们报告了该方法的计算效率和统计可靠性,并用DNN进行了非分布检测。
Uncertainty evaluation is a core technique when deep neural networks (DNNs) are used in real-world problems. In practical applications, we often encounter unexpected samples that have not seen in the training process. Not only achieving the high-prediction accuracy but also detecting uncertain data is significant for safety-critical systems. In statistics and machine learning, Bayesian inference has been exploited for uncertainty evaluation. The Bayesian neural networks (BNNs) have recently attracted considerable attention in this context, as the DNN trained using dropout is interpreted as a Bayesian method. Based on this interpretation, several methods to calculate the Bayes predictive distribution for DNNs have been developed. Though the Monte-Carlo method called MC dropout is a popular method for uncertainty evaluation, it requires a number of repeated feed-forward calculations of DNNs with randomly sampled weight parameters. To overcome the computational issue, we propose a sampling-free method to evaluate uncertainty. Our method converts a neural network trained using dropout to the corresponding Bayesian neural network withvariance propagation. Our method is available not only to feed-forward NNs but also to recurrent NNs such as LSTM. We report the computational efficiency and statistical reliability of our method in numerical experiments of language modeling using RNNs, and the out-of-distribution detection with DNNs.