A Tail-Index Analysis of Stochastic Gradient Noise in Deep Neural Networks

A Tail-Index Analysis of Stochastic Gradient Noise in Deep Neural Networks
复制标题

DOI:
--
复制
发表时间:
2019-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Umut Simsekli;Levent Sagun;M. Gürbüzbalaban
Umut Simsekli;Levent Sagun;M. Gürbüzbalaban
中科院分区:
其他
文献类型:
--
作者:
Umut Simsekli;Levent Sagun;M. Gürbüzbalaban

文献摘要

被引文献

相似文献

随机梯度下降(SGD)算法中的梯度噪声(GN)在大数据量情况下通常被认为是高斯的,这是因为假设经典的中心极限定理(CLT)起作用。这一假设通常是为了数学上的方便,因为它使SGD能够被分析为由布朗运动驱动的随机微分方程(SDE)。我们认为,高斯假设在深度学习环境中可能不成立,从而使基于布朗运动的分析变得不合适。受非高斯自然现象的启发,我们在更一般的背景下考虑广义随机变量,并引用广义CLT(GCLT),它表明广义随机变量收敛于重尾$\α$稳定的随机变量。因此,我们建议将SGD分析为由Levy动议驱动的SDE。正如现有的亚稳理论所证明的那样,这种SDE可能会引起“跳跃”,迫使SDE从狭窄的极小值过渡到更宽的极小值。为了验证$\α$稳定的假设,我们在常见的深度学习结构上进行了大量的实验,结果表明,在所有的设置下,GN都是高度非高斯性的,并且允许重尾。我们进一步研究了在不同的网络架构和大小、损失函数和数据集下的尾部行为。我们的结果打开了一个不同的视角,并更多地阐明了SGD更喜欢广义极小值的信念。
The gradient noise (GN) in the stochastic gradient descent (SGD) algorithm is often considered to be Gaussian in the large data regime by assuming that the classical central limit theorem (CLT) kicks in. This assumption is often made for mathematical convenience, since it enables SGD to be analyzed as a stochastic differential equation (SDE) driven by a Brownian motion. We argue that the Gaussianity assumption might fail to hold in deep learning settings and hence render the Brownian motion-based analyses inappropriate. Inspired by non-Gaussian natural phenomena, we consider the GN in a more general context and invoke the generalized CLT (GCLT), which suggests that the GN converges to a heavy-tailed $\alpha$-stable random variable. Accordingly, we propose to analyze SGD as an SDE driven by a Levy motion. Such SDEs can incur `jumps', which force the SDE transition from narrow minima to wider minima, as proven by existing metastability theory. To validate the $\alpha$-stable assumption, we conduct extensive experiments on common deep learning architectures and show that in all settings, the GN is highly non-Gaussian and admits heavy-tails. We further investigate the tail behavior in varying network architectures and sizes, loss functions, and datasets. Our results open up a different perspective and shed more light on the belief that SGD prefers wide minima.