The Dropout Learning Algorithm.

The Dropout Learning Algorithm.
复制标题

DOI:
10.1016/j.artint.2014.02.004
复制
发表时间:
2014-05
影响因子:
14.4
通讯作者:
Sadowski, Peter
Sadowski, Peter
中科院分区:
计算机科学2区
文献类型:
--
作者:
Baldi, Pierre;Sadowski, Peter

文献摘要

参考文献

被引文献

相似文献

Dropout是最近引入的一种用于训练神经网络的算法,通过在训练过程中随机丢弃单元来防止它们的共同适应。使用伯努利门控变量,一般足以适应单元或连接上的辍学,并与可变速率的一些静态和动态特性的辍学的数学分析。该框架允许完整的分析线性网络中dropout的系综平均特性,这对于理解非线性情况是有用的。非线性Logistic网络中dropout的系综平均性质是由三个基本方程得到的:(1)Logistic函数的期望值用归一化几何平均逼近,并给出了其界和估计:(2)Logistic函数的归一化几何平均与均值的Logistic之间的代数等式,它在数学上刻画了Logistic函数;以及(3)平均值相对于自变量的和以及乘积的线性。结果也扩展到其他类别的传递函数,包括整流线性函数。近似误差往往相互抵消,不会累积。Dropout也可以连接到随机神经元并用于预测发射率,以及通过将反向传播视为Dropout线性网络中的整体平均来连接到反向传播。此外,dropout的收敛性质可以用随机梯度下降来理解。最后,对于dropout的正则化性质,dropout梯度的期望值是相应近似集合的梯度,由具有自洽方差最小化和稀疏表示倾向的自适应权重衰减项正则化。
Dropout is a recently introduced algorithm for training neural network by randomly dropping units during training to prevent their co-adaptation. A mathematical analysis of some of the static and dynamic properties of dropout is provided using Bernoulli gating variables, general enough to accommodate dropout on units or connections, and with variable rates. The framework allows a complete analysis of the ensemble averaging properties of dropout in linear networks, which is useful to understand the non-linear case. The ensemble averaging properties of dropout in non-linear logistic networks result from three fundamental equations: (1) the approximation of the expectations of logistic functions by normalized geometric means, for which bounds and estimates are derived; (2) the algebraic equality between normalized geometric means of logistic functions with the logistic of the means, which mathematically characterizes logistic functions; and (3) the linearity of the means with respect to sums, as well as products of independent variables. The results are also extended to other classes of transfer functions, including rectified linear functions. Approximation errors tend to cancel each other and do not accumulate. Dropout can also be connected to stochastic neurons and used to predict firing rates, and to backpropagation by viewing the backward propagation as ensemble averaging in a dropout linear network. Moreover, the convergence properties of dropout can be understood in terms of stochastic gradient descent. Finally, for the regularization properties of dropout, the expectation of the dropout gradient is the gradient of the corresponding approximation ensemble, regularized by an adaptive weight decay term with a propensity for self-consistent variance minimization and sparse representations.
DOI: 10.1016/0893-6080(89)90014-2
发表时间: 1989-01-01
期刊: NEURAL NETWORKS
影响因子: 7.8
作者:
BALDI, P;HORNIK, K
通讯作者: HORNIK, K
DOI: 10.1162/neco.1996.8.3.643
发表时间: 1996-04-01
期刊: NEURAL COMPUTATION
影响因子: 2.9
作者:
An, GZ
通讯作者: An, GZ
DOI: 10.3926/jiem.2009.v2n1.p114-127
发表时间: 2009-01-01
影响因子: 3
作者:
Bowling, Shannon R.;Khasawneh, Mohammad T.;Cho, Byung Rae
通讯作者: Cho, Byung Rae
DOI: 10.1007/bf00058655
发表时间: 1996-08-01
期刊: MACHINE LEARNING
影响因子: 7.5
作者:
Breiman, L
通讯作者: Breiman, L
DOI: 10.1016/0893-6080(89)90016-6
发表时间: 1989-01-01
期刊: NEURAL NETWORKS
影响因子: 7.8
作者:
GARDNER, D
通讯作者: GARDNER, D