On the Global Convergence of Training Deep Linear ResNets

On the Global Convergence of Training Deep Linear ResNets
复制标题

DOI:
--
复制
发表时间:
2020-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Difan Zou;Philip M. Long;Quanquan Gu
Difan Zou;Philip M. Long;Quanquan Gu
中科院分区:
其他
文献类型:
--
作者:
Difan Zou;Philip M. Long;Quanquan Gu

文献摘要

相似文献

本文研究了梯度下降(GD)和随机梯度下降(SGD)训练L$-隐层线性残差网络(ResNets)的收敛性。我们证明了,对于在输入和输出层具有一定线性变换的深度残差网络(在整个训练过程中是固定的),在所有隐藏权重上具有零初始化的GD和SGD都可以收敛到训练损失的全局最小值。此外,当专门针对适当的高斯随机线性变换时,GD和SGD可证明优化了足够宽的深度线性ResNet。与GD训练标准深度线性网络的全局收敛结果(Du & Hu 2019)相比,我们对神经网络宽度的条件更尖锐,是$O(\kappa L)$的一个因子,其中$\kappa$表示训练数据协方差矩阵的条件数。我们进一步提出了一种改进的恒等输入输出变换,并证明了$(d+k)$-wide神经网络足以保证GD/SGD的全局收敛性,其中$d,k$分别是输入和输出维数.
We study the convergence of gradient descent (GD) and stochastic gradient descent (SGD) for training $L$-hidden-layer linear residual networks (ResNets). We prove that for training deep residual networks with certain linear transformations at input and output layers, which are fixed throughout training, both GD and SGD with zero initialization on all hidden weights can converge to the global minimum of the training loss. Moreover, when specializing to appropriate Gaussian random linear transformations, GD and SGD provably optimize wide enough deep linear ResNets. Compared with the global convergence result of GD for training standard deep linear networks (Du & Hu 2019), our condition on the neural network width is sharper by a factor of $O(\kappa L)$, where $\kappa$ denotes the condition number of the covariance matrix of the training data. We further propose a modified identity input and output transformations, and show that a $(d+k)$-wide neural network is sufficient to guarantee the global convergence of GD/SGD, where $d,k$ are the input and output dimensions respectively.