Information Geometry of Orthogonal Initializations and Training

Information Geometry of Orthogonal Initializations and Training
复制标题

DOI:
--
复制
发表时间:
2018-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Piotr A. Sokól;Il-Su Park
Piotr A. Sokól;Il-Su Park
中科院分区:
其他
文献类型:
--
作者:
Piotr A. Sokól;Il-Su Park

文献摘要

被引文献

相似文献

最近,平均场理论已成功地用于分析宽随机神经网络的性质。它产生了一个规范性的理论,用于初始化具有正交权重的前馈神经网络,这确保了正向传播的激活和反向传播的梯度都接近等距,因此训练速度快了几个数量级。尽管有很强的经验性能,但人们对关键初始化在深度神经网络优化中赋予优势的机制知之甚少。在这里,我们展示了由Fisher信息矩阵(Fisher information matrix)测量的优化景观的最大曲率(梯度平滑度)与输入输出雅可比矩阵的谱半径之间的一种新联系,这部分解释了为什么更多的等距网络可以更快地训练。此外,考虑到正交权重对于确保梯度范数在初始化时近似保持是必要的,我们通过实验研究了在整个训练过程中保持正交性的好处,从中我们得出结论,无论梯度的平滑度如何,权重的流形优化都表现良好。此外,由实验结果的动机,我们表明,一个低的条件数的学习速度更快的预测。
Recently mean field theory has been successfully used to analyze properties of wide, random neural networks. It gave rise to a prescriptive theory for initializing feed-forward neural networks with orthogonal weights, which ensures that both the forward propagated activations and the backpropagated gradients are near $\ell_2$ isometries and as a consequence training is orders of magnitude faster. Despite strong empirical performance, the mechanisms by which critical initializations confer an advantage in the optimization of deep neural networks are poorly understood. Here we show a novel connection between the maximum curvature of the optimization landscape (gradient smoothness) as measured by the Fisher information matrix (FIM) and the spectral radius of the input-output Jacobian, which partially explains why more isometric networks can train much faster. Furthermore, given that orthogonal weights are necessary to ensure that gradient norms are approximately preserved at initialization, we experimentally investigate the benefits of maintaining orthogonality throughout training, from which we conclude that manifold optimization of weights performs well regardless of the smoothness of the gradients. Moreover, motivated by experimental results we show that a low condition number of the FIM is not predictive of faster learning.