Data Augmentation for Bayesian Deep Learning

Data Augmentation for Bayesian Deep Learning
复制标题

贝叶斯深度学习的数据增强

DOI:
10.1214/22-ba1331
复制
发表时间:
2019-03
期刊:
影响因子:
4.4
通讯作者:
YueXing Wang;Nicholas G. Polson;Vadim O. Sokolov
YueXing Wang;Nicholas G. Polson;Vadim O. Sokolov
中科院分区:
数学2区
文献类型:
--
作者:
YueXing Wang;Nicholas G. Polson;Vadim O. Sokolov

文献摘要

相似文献

深度学习(DL)方法已成为函数逼近和预测的最强大工具之一。虽然DL的表示属性已经得到了很好的研究,但不确定性量化仍然具有挑战性,并且在很大程度上尚未探索。数据增强技术是一种自然的方法,以提供不确定性量化,并将随机蒙特卡罗搜索到随机梯度下降(SGD)方法。我们的论文的目的是表明,训练DL架构与数据增强导致效率的提高。我们使用法线的尺度混合理论来推导深度学习的数据增强策略。这使得期望最大化和MCMC算法的变体可以应用于这些高维非线性深度学习模型。为了展示我们的方法,我们为各种常用的激活函数开发了数据增强算法:logit,ReLU,leaky ReLU和SVM。我们的方法相比,传统的随机梯度下降与反向传播。我们的优化过程导致一个版本的迭代重新加权最小二乘法,并可以实现规模与加速线性代数方法提供了显着的速度提高。我们在一些标准数据集上说明了我们的方法。最后,我们总结了未来研究的方向。
Deep Learning (DL) methods have emerged as one of the most powerful tools for functional approximation and prediction. While the representation properties of DL have been well studied, uncertainty quantification remains challenging and largely unexplored. Data augmentation techniques are a natural approach to provide uncertainty quantification and to incorporate stochastic Monte Carlo search into stochastic gradient descent (SGD) methods. The purpose of our paper is to show that training DL architectures with data augmentation leads to efficiency gains. We use the theory of scale mixtures of normals to derive data augmentation strategies for deep learning. This allows variants of the expectation-maximization and MCMC algorithms to be brought to bear on these high dimensional nonlinear deep learning models. To demonstrate our methodology, we develop data augmentation algorithms for a variety of commonly used activation functions: logit, ReLU, leaky ReLU and SVM. Our methodology is compared to traditional stochastic gradient descent with back-propagation. Our optimization procedure leads to a version of iteratively re-weighted least squares and can be implemented at scale with accelerated linear algebra methods providing substantial improvement in speed. We illustrate our methodology on a number of standard datasets. Finally, we conclude with directions for future research.