Joint Inference for Neural Network Depth and Dropout Regularization

Joint Inference for Neural Network Depth and Dropout Regularization
复制标题

DOI:
--
复制
发表时间:
2021
期刊:
--
影响因子:
--
通讯作者:
C. KishanK.;Rui Li;MohammadMahdi Gilany
C. KishanK.;Rui Li;MohammadMahdi Gilany
中科院分区:
其他
文献类型:
--
作者:
C. KishanK.;Rui Li;MohammadMahdi Gilany

文献摘要

相似文献

Dropout正则化方法修剪神经网络的预定骨干结构,以避免过度拟合。然而,深度模型仍然倾向于校准不良,对不正确的预测具有很高的信心。我们提出了一种统一的艾德贝叶斯模型选择方法,以联合推断数据所保证的最合理的网络深度,并同时执行dropout正则化。特别是,为了推断网络深度,我们定义了一个beta过程,这个过程覆盖了隐藏层的数量,允许它到达无穷大。由β过程诱导的逐层激活概率通过共轭伯努利过程的二进制向量来调节神经元激活。跨域的实验表明,通过调整网络深度和辍学正则化数据,我们的方法实现了上级性能相比,最先进的方法校准良好的不确定性估计。在持续学习中,我们的方法使神经网络能够动态地发展其深度,以适应超出其初始结构的增量可用数据,并减轻灾难性遗忘。
Dropout regularization methods prune a neural network’s pre-determined backbone structure to avoid overfitting. However, a deep model still tends to be poorly calibrated with high confidence on incorrect predictions. We propose a unified Bayesian model selection method to jointly infer the most plausible network depth warranted by data, and perform dropout regularization simultaneously. In particular, to infer network depth we define a beta process over the number of hidden layers which allows it to go to infinity. Layer-wise activation probabilities induced by the beta process modulate neuron activation via binary vectors of a conjugate Bernoulli process. Experiments across domains show that by adapting network depth and dropout regularization to data, our method achieves superior performance comparing to state-of-the-art methods with well-calibrated uncertainty estimates. In continual learning, our method enables neural networks to dynamically evolve their depths to accommodate incrementally available data beyond their initial structures, and alleviate catastrophic forgetting.