Bad Global Minima Exist and SGD Can Reach Them

Bad Global Minima Exist and SGD Can Reach Them
复制标题

DOI:
--
复制
发表时间:
2019-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Shengchao Liu;Dimitris Papailiopoulos;D. Achlioptas
Shengchao Liu;Dimitris Papailiopoulos;D. Achlioptas
中科院分区:
其他
文献类型:
--
作者:
Shengchao Liu;Dimitris Papailiopoulos;D. Achlioptas

文献摘要

被引文献

相似文献

最近的几项工作旨在解释为什么严重过度参数化的模型在随机梯度下降(SGD)训练时泛化良好。紧急共识解释有两个部分:第一个是“没有坏的局部极小值”,第二个是SGD通过偏向低复杂度模型来执行隐式正则化。我们在使用常见的深度神经网络架构进行图像分类的背景下重新审视了这两个想法。我们的第一个发现是存在坏的全局极小值,即,模型完美地拟合训练集,但泛化能力差。我们的第二个发现是,只给未标记的训练数据,我们可以很容易地构造初始化,这将导致SGD快速收敛到这种糟糕的全局最小值。例如,在CIFAR、CINIC10和(受限)ImageNet上,这可以通过在训练数据上拟合随机标签得到的模型上开始SGD来实现:虽然后续的SGD训练(使用正确的标签)将达到零训练误差,但与随机初始化的训练相比,最终模型的测试精度将下降高达40%。最后,我们证明了正则化似乎为SGD提供了一条逃生路线:一旦使用了诸如数据增强之类的竞争,从复杂模型(对抗初始化)开始对测试精度没有影响。
Several recent works have aimed to explain why severely overparameterized models, generalize well when trained by Stochastic Gradient Descent (SGD). The emergent consensus explanation has two parts: the first is that there are "no bad local minima", while the second is that SGD performs implicit regularization by having a bias towards low complexity models. We revisit both of these ideas in the context of image classification with common deep neural network architectures. Our first finding is that there exist bad global minima, i.e., models that fit the training set perfectly, yet have poor generalization. Our second finding is that given only unlabeled training data, we can easily construct initializations that will cause SGD to quickly converge to such bad global minima. For example, on CIFAR, CINIC10, and (Restricted) ImageNet, this can be achieved by starting SGD at a model derived by fitting random labels on the training data: while subsequent SGD training (with the correct labels) will reach zero training error, the resulting model will exhibit a test accuracy degradation of up to 40% compared to training from a random initialization. Finally, we show that regularization seems to provide SGD with an escape route: once heuristics such as data augmentation are used, starting from a complex model (adversarial initialization) has no effect on the test accuracy.