Algorithm-Dependent Generalization Bounds for Overparameterized Deep Residual Networks

Algorithm-Dependent Generalization Bounds for Overparameterized Deep Residual Networks
复制标题

DOI:
--
复制
发表时间:
2019-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Spencer Frei;Yuan Cao;Quanquan Gu
Spencer Frei;Yuan Cao;Quanquan Gu
中科院分区:
其他
文献类型:
--
作者:
Spencer Frei;Yuan Cao;Quanquan Gu

文献摘要

相似文献

由于具有这种架构的网络的泛化性和稳定性的提高,残差网络中使用的跳跃连接已成为深度学习中的标准架构选择,尽管这种改进的性能的理论保证有限。在这项工作中,我们分析了随机初始化后通过梯度下降训练的过参数化深度残差网络,并证明(i)通过梯度下降学习的网络类构成了整个神经网络函数类的一个小子集,并且(ii)这个网络子类足够大以保证较小的训练误差。通过显示(i),我们能够证明使用梯度下降训练的深度残差网络在训练和测试误差之间具有较小的泛化差距,并且与(ii)一起保证了测试误差很小。我们的优化和泛化保证需要在网络深度上仅是对数的超参数化,这有助于解释为什么残差网络比全连接网络更可取。
The skip-connections used in residual networks have become a standard architecture choice in deep learning due to the increased generalization and stability of networks with this architecture, although there have been limited theoretical guarantees for this improved performance. In this work, we analyze overparameterized deep residual networks trained by gradient descent following random initialization, and demonstrate that (i) the class of networks learned by gradient descent constitutes a small subset of the entire neural network function class, and (ii) this subclass of networks is sufficiently large to guarantee small training error. By showing (i) we are able to demonstrate that deep residual networks trained with gradient descent have a small generalization gap between training and test error, and together with (ii) this guarantees that the test error will be small. Our optimization and generalization guarantees require overparameterization that is only logarithmic in the depth of the network, which helps explain why residual networks are preferable to fully connected ones.