Jumpout : Improved Dropout for Deep Neural Networks with ReLUs

Jumpout : Improved Dropout for Deep Neural Networks with ReLUs
复制标题

DOI:
--
复制
发表时间:
2019-05
期刊:
--
影响因子:
--
通讯作者:
Shengjie Wang;Tianyi Zhou;J. Bilmes
Shengjie Wang;Tianyi Zhou;J. Bilmes
中科院分区:
其他
文献类型:
--
作者:
Shengjie Wang;Tianyi Zhou;J. Bilmes

文献摘要

被引文献

相似文献

我们讨论了关于带ReLUs的DNN的dropout的三个新见解:1)dropout鼓励DNN的每个局部线性块在附近区域的数据点上进行训练;2)相同的dropout率导致具有不同部分的ReLUdeactivated神经元的层的不同(有效)失活率;3) dropout的重标度因子与批归一化一起使用时,会导致训练和测试之间的归一化不一致。上面的代码导致了三个简单但重要的修改,从而导致我们的方法“跳出”。Jumpout对单调递减分布(例如,高斯分布的右半部分)的辍学率进行采样,因此每个局部线性块都被高概率地训练,以更好地处理来自附近而不是更远区域的数据点。此外,Jumpout对每一层和每一批训练的退出率进行自适应归一化,使激活神经元的有效失活率保持不变。此外,它重新调整输出以获得更好的权衡,使神经元的方差和均值在训练和测试阶段之间更加一致,从而减轻dropout和批处理归一化之间的不兼容性。Jumpout显著提高了CIFAR10、CIFAR100、Fashion-MNIST、STL10、SVHN、ImageNet-1k等不同神经网络的性能,而引入的额外内存和计算成本可以忽略不计。
We discuss three novel insights about dropout for DNNs with ReLUs: 1) dropout encourages each local linear piece of a DNN to be trained on data points from nearby regions; 2) the same dropout rate results in different (effective) deactivation rates for layers with different portions of ReLUdeactivated neurons; and 3) the rescaling factor of dropout causes a normalization inconsistency between training and test when used together with batch normalization. The above leads to three simple but nontrivial modifications resulting in our method “jumpout.” Jumpout samples the dropout rate from a monotone decreasing distribution (e.g., the right half of a Gaussian), so each local linear piece is trained, with high probability, to work better for data points from nearby than more distant regions. Jumpout moreover adaptively normalizes the dropout rate at each layer and every training batch, so the effective deactivation rate on the activated neurons is kept the same. Furthermore, it rescales the outputs for a better trade-off that keeps both the variance and mean of neurons more consistent between training and test phases, thereby mitigating the incompatibility between dropout and batch normalization. Jumpout significantly improves the performance of different neural nets on CIFAR10, CIFAR100, Fashion-MNIST, STL10, SVHN, ImageNet-1k, etc., while introducing negligible additional memory and computation costs.