Equivalence between dropout and data augmentation: A mathematical check

Equivalence between dropout and data augmentation: A mathematical check
复制标题

DOI:
10.1016/j.neunet.2019.03.013
复制
发表时间:
2019-07
期刊:
Neural networks : the official journal of the International Neural Network Society
影响因子:
--
通讯作者:
Dazhi Zhao;Guozhu Yu;Peng Xu;M. Luo
Dazhi Zhao;Guozhu Yu;Peng Xu;M. Luo
中科院分区:
其他
文献类型:
--
作者:
Dazhi Zhao;Guozhu Yu;Peng Xu;M. Luo

文献摘要

被引文献

相似文献

深度学习的巨大成就可以归功于其强大的特征表示能力,其中表示能力来自于非线性激活函数和大量的网络节点。然而,深度神经网络存在收敛速度慢等严重问题,丢包是提高网络泛化能力和测试性能的一种突出方法。关于辍学效果如此之好的原因,人们给出了许多解释,其中,辍学与数据扩充之间的等价性是一种新提出的、具有启发性的解释。在本文中,我们讨论了这种等价性成立的确切条件。我们的主要结果保证,当输入空间的维度等于或高于输出空间的维度时,等价关系几乎必然成立。此外,如果将常用的修正后的线性单元激活函数替换为新提出的值位于R中的激活函数,则我们的结果可以推广到多层神经网络。为便于比较,文中还给出了不等价情形的反例。最后,在MNIST数据集上进行了一系列实验,以说明和帮助理解理论结果。
The great achievements of deep learning can be attributed to its tremendous power of feature representation, where the representation ability comes from the nonlinear activation function and the large number of network nodes. However, deep neural networks suffer from serious issues such as slow convergence, and dropout is an outstanding method to improve the network’s generalization ability and test performance. Many explanations have been given for why dropout works so well, among which the equivalence between dropout and data augmentation is a newly proposed and stimulating explanation. In this article, we discuss the exact conditions for this equivalence to hold. Our main result guarantees that the equivalence relation almost surely holds if the dimension of the input space is equal to or higher than that of the output space. Furthermore, if the commonly used rectified linear unit activation function is replaced by some newly proposed activation function whose value lies in R, then our results can be extended to multilayer neural networks. For comparison, some counterexamples are given for the inequivalent case. Finally, a series of experiments on the MNIST dataset are conducted to illustrate and help understand the theoretical results.