Understanding the Generalization of Adam in Learning Neural Networks with Proper Regularization

Understanding the Generalization of Adam in Learning Neural Networks with Proper Regularization
复制标题

DOI:
--
复制
发表时间:
2021-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Difan Zou;Yuan Cao;Yuanzhi Li;Quanquan Gu
Difan Zou;Yuan Cao;Yuanzhi Li;Quanquan Gu
中科院分区:
其他
文献类型:
--
作者:
Difan Zou;Yuan Cao;Yuanzhi Li;Quanquan Gu

文献摘要

被引文献

相似文献

ADAM等自适应梯度算法在深度学习优化中得到了越来越广泛的应用。然而,已经观察到,与(随机)梯度下降相比,在许多深度学习应用中,即使在微调的正则化的情况下,ADAM也可以收敛到不同的解,并且测试误差明显更差。本文给出了这一现象的理论解释:我们证明了在从相同的随机初始化开始学习的过参数两层卷积神经网络的非凸设置下,对于一类数据分布(来自图像数据),ADAM和梯度下降(GD)算法可以收敛到训练目标的不同全局解,且具有明显不同的泛化误差,即使采用权重衰减正则化。相反,我们证明了如果训练目标是凸的,并且采用了权重衰减正则化,则任何优化算法,包括ADAM和GD,如果训练成功,都会收敛到相同的解。这表明ADAM较差的泛化性能从根本上与深度学习优化的非凸性环境有关。
Adaptive gradient methods such as Adam have gained increasing popularity in deep learning optimization. However, it has been observed that compared with (stochastic) gradient descent, Adam can converge to a different solution with a significantly worse test error in many deep learning applications such as image classification, even with a fine-tuned regularization. In this paper, we provide a theoretical explanation for this phenomenon: we show that in the nonconvex setting of learning over-parameterized two-layer convolutional neural networks starting from the same random initialization, for a class of data distributions (inspired from image data), Adam and gradient descent (GD) can converge to different global solutions of the training objective with provably different generalization errors, even with weight decay regularization. In contrast, we show that if the training objective is convex, and the weight decay regularization is employed, any optimization algorithms including Adam and GD will converge to the same solution if the training is successful. This suggests that the inferior generalization performance of Adam is fundamentally tied to the nonconvex landscape of deep learning optimization.