Entropy-SGD: biasing gradient descent into wide valleys

Entropy-SGD: biasing gradient descent into wide valleys
复制标题

DOI:
10.1088/1742-5468/ab39d9
复制
发表时间:
2019-12-01
影响因子:
2.4
通讯作者:
Zecchina, Riccardo
Zecchina, Riccardo
中科院分区:
物理与天体物理3区
文献类型:
--
作者:
Chaudhari, Pratik;Choromanska, Anna;Zecchina, Riccardo

文献摘要

被引文献

相似文献

本文提出了一种名为Entropy-SGD的新优化算法,用于训练深度神经网络,该算法由能量景观的局部几何结构驱动。具有低推广误差的局部极值在Hessian中具有很大比例的几乎为零的特征值,具有很少的正或负特征值。我们利用这一观察,构建一个局部熵为基础的目标函数,有利于良好的概括性的解决方案,躺在大的平坦地区的能源景观,同时避免较差的概括性的解决方案位于尖锐的山谷。从概念上讲,我们的算法类似于SGD的两个嵌套循环,其中我们在内部循环中使用Langevin动力学来计算每次更新权重之前的局部熵的梯度。我们表明,新的目标具有更平滑的能量景观,并在一定的假设下,使用一致稳定性,显示出比SGD更好的泛化能力。我们在卷积和递归网络上的实验表明,熵-SGD在泛化误差和训练时间方面优于最先进的技术。
This paper proposes a new optimization algorithm called Entropy-SGD for training deep neural networks that is motivated by the local geometry of the energy landscape. Local extrema with low generalization error have a large proportion of almost-zero eigenvalues in the Hessian with very few positive or negative eigenvalues. We leverage upon this observation to construct a local-entropy-based objective function that favors well-generalizable solutions lying in large flat regions of the energy landscape, while avoiding poorly-generalizable solutions located in the sharp valleys. Conceptually, our algorithm resembles two nested loops of SGD where we use Langevin dynamics in the inner loop to compute the gradient of the local entropy before each update of the weights. We show that the new objective has a smoother energy landscape and show improved generalization over SGD using uniform stability, under certain assumptions. Our experiments on convolutional and recurrent networks demonstrate that Entropy-SGD compares favorably to state-of-the-art techniques in terms of generalization error and training time.