Shakedrop Regularization for Deep Residual Learning

Shakedrop Regularization for Deep Residual Learning
复制标题

DOI:
10.1109/access.2019.2960566
复制
发表时间:
2018-02
期刊:
影响因子:
3.9
通讯作者:
Yoshihiro Yamada;M. Iwamura;Takuya Akiba;K. Kise
Yoshihiro Yamada;M. Iwamura;Takuya Akiba;K. Kise
中科院分区:
计算机科学3区
文献类型:
--
作者:
Yoshihiro Yamada;M. Iwamura;Takuya Akiba;K. Kise

文献摘要

被引文献

相似文献

过拟合是深度神经网络中的一个关键问题,即使在最新的网络结构中也是如此。本文针对ResNet及其改进算法(如宽ResNet、金字塔网和ResNeXt)的过拟合问题,提出了一种新的正则化方法ShakeDrop正则化方法。ShakeDrop的灵感来自Shake-Shake,这是一种有效的正则化方法,但仅适用于ResNeXt。ShakeDrop比Shake-Shake更有效,不仅可以应用于ResNeXt,还可以应用于ResNet、Wide ResNet和金字塔网。一个重要的关键是实现训练的稳定性。由于有效的正则化往往会导致训练不稳定,我们引入了训练稳定器,这是现有正则化方法的一种不寻常的使用。通过在不同条件下的实验,论证了ShakeDrop在不同条件下工作良好的条件。
Overfitting is a crucial problem in deep neural networks, even in the latest network architectures. In this paper, to relieve the overfitting effect of ResNet and its improvements (i.e., Wide ResNet, PyramidNet, and ResNeXt), we propose a new regularization method called ShakeDrop regularization. ShakeDrop is inspired by Shake-Shake, which is an effective regularization method, but can be applied to ResNeXt only. ShakeDrop is more effective than Shake-Shake and can be applied not only to ResNeXt but also ResNet, Wide ResNet, and PyramidNet. An important key is to achieve stability of training. Because effective regularization often causes unstable training, we introduce a training stabilizer, which is an unusual use of an existing regularizer. Through experiments under various conditions, we demonstrate the conditions under which ShakeDrop works well.