On Generalization Bounds for Deep Networks Based on Loss Surface Implicit Regularization

On Generalization Bounds for Deep Networks Based on Loss Surface Implicit Regularization
复制标题

DOI:
10.1109/tit.2022.3215088
复制
发表时间:
2022-01
影响因子:
2.5
通讯作者:
M. Imaizumi;J. Schmidt-Hieber
M. Imaizumi;J. Schmidt-Hieber
中科院分区:
计算机科学2区
文献类型:
--
作者:
M. Imaizumi;J. Schmidt-Hieber

文献摘要

相似文献

经典的统计学习理论认为,拟合过多的参数会导致过拟合和性能差。尽管存在大量参数,但现代深度神经网络仍能很好地泛化,这与这一发现相矛盾,并构成了解释深度学习成功的主要未解决问题。以往的工作主要集中在随机梯度下降(SGD)引起的隐式正则化,而我们在这里研究了在高斯梯度噪声下,局部极小值周围的能量景观的局部几何形状如何影响SGD的统计特性。我们认为,在合理的假设下,局部几何迫使SGD保持接近低维子空间,这导致了另一种形式的隐式正则化,并导致深度神经网络泛化误差的更严格界限。为了推导神经网络的泛化误差边界,我们首先在局部极小值周围引入停滞集的概念,并施加种群风险的局部本质凸性。在这些条件下,导出了SGD保持在这些停滞集中的下界。如果出现停滞,我们推导了深度神经网络泛化误差的界,涉及权矩阵的谱范数,而不是网络参数的数量。从技术上讲,我们的证明是基于控制SGD迭代中参数值的变化和基于局部最小值周围合适邻域的熵的经验损失函数的局部一致收敛。
The classical statistical learning theory implies that fitting too many parameters leads to overfitting and poor performance. That modern deep neural networks generalize well despite a large number of parameters contradicts this finding and constitutes a major unsolved problem towards explaining the success of deep learning. While previous work focuses on the implicit regularization induced by stochastic gradient descent (SGD), we study here how the local geometry of the energy landscape around local minima affects the statistical properties of SGD with Gaussian gradient noise. We argue that under reasonable assumptions, the local geometry forces SGD to stay close to a low dimensional subspace and that this induces another form of implicit regularization and results in tighter bounds on the generalization error for deep neural networks. To derive generalization error bounds for neural networks, we first introduce a notion of stagnation sets around the local minima and impose a local essential convexity property of the population risk. Under these conditions, lower bounds for SGD to remain in these stagnation sets are derived. If stagnation occurs, we derive a bound on the generalization error of deep neural networks involving the spectral norms of the weight matrices but not the number of network parameters. Technically, our proofs are based on controlling the change of parameter values in the SGD iterates and local uniform convergence of the empirical loss functions based on the entropy of suitable neighborhoods around local minima.