A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima

A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima
复制标题

DOI:
--
复制
发表时间:
2020-02
期刊:
arXiv: Learning
影响因子:
--
通讯作者:
Zeke Xie;Issei Sato;Masashi Sugiyama
Zeke Xie;Issei Sato;Masashi Sugiyama
中科院分区:
其他
文献类型:
--
作者:
Zeke Xie;Issei Sato;Masashi Sugiyama

文献摘要

相似文献

随机梯度下降(SGD)及其变体是实践中训练深度网络的主流方法。众所周知,SGD 可以找到一个通常可以很好概括的平坦最小值。然而,从数学上讲,尚不清楚深度学习如何在如此众多的最小值中选择一个平坦的最小值。为了定量地回答这个问题,我们发展了密度扩散理论(DDT)来揭示最小选择如何定量地依赖于最小清晰度和超参数。据我们所知,我们是第一个从理论上和经验上证明,受益于随机梯度噪声的 Hessian 相关协方差,SGD 对平坦最小值的支持指数级高于锐最小值,而注入白噪声的梯度下降 (GD) 仅在多项式上比锐最小值更支持平坦最小值。我们还发现,小学习率或大批量训练都需要指数级多次迭代才能摆脱批量大小与学习率之比的最小值。因此,大批量训练无法在实际计算时间内有效地搜索平坦最小值。
Stochastic Gradient Descent (SGD) and its variants are mainstream methods for training deep networks in practice. SGD is known to find a flat minimum that often generalizes well. However, it is mathematically unclear how deep learning can select a flat minimum among so many minima. To answer the question quantitatively, we develop a density diffusion theory (DDT) to reveal how minima selection quantitatively depends on the minima sharpness and the hyperparameters. To the best of our knowledge, we are the first to theoretically and empirically prove that, benefited from the Hessian-dependent covariance of stochastic gradient noise, SGD favors flat minima exponentially more than sharp minima, while Gradient Descent (GD) with injected white noise favors flat minima only polynomially more than sharp minima. We also reveal that either a small learning rate or large-batch training requires exponentially many iterations to escape from minima in terms of the ratio of the batch size and learning rate. Thus, large-batch training cannot search flat minima efficiently in a realistic computational time.