Understanding approximate Fisher information for fast convergence of natural gradient descent in wide neural networks*

Understanding approximate Fisher information for fast convergence of natural gradient descent in wide neural networks*
复制标题

了解宽神经网络中自然梯度下降快速收敛的近似 Fisher 信息*

DOI:
10.1088/1742-5468/ac3ae3
复制
发表时间:
2021
期刊:
Journal of Statistical Mechanics: Theory and Experiment
影响因子:
--
通讯作者:
Osawa Kazuki
Osawa Kazuki
中科院分区:
--
文献类型:
--
作者:
Karakida Ryo;Osawa Kazuki

文献摘要

相似文献

自然梯度下降(NGD)有助于加速梯度下降动力学的收敛,但由于计算成本较高,需要在大规模深度神经网络中进行近似。实证研究证实,一些具有近似Fisher信息的NGD方法在实践中收敛得足够快。然而,从理论角度来看,仍然不清楚为什么以及在什么条件下这种启发式近似有效。在这项工作中,我们揭示了在特定条件下,具有近似 Fisher 信息的 NGD 可以实现与精确 NGD 相同的快速收敛到全局最小值。我们考虑无限宽度限制下的深度神经网络,并通过神经正切核分析函数空间中 NGD 的渐近训练动力学。在函数空间中,具有近似Fisher信息的训练动力学与具有精确Fisher信息的训练动力学相同,并且它们收敛得很快。快速收敛在逐层近似中成立;例如,在块对角近似中,每个块对应于一个层,以及在块三对角近似和 K-FAC 近似中。我们还发现,在某些假设下,单位近似可以实现相同的快速收敛。所有这些不同的近似在函数空间中都具有各向同性梯度,这对于在训练中实现相同的收敛特性起着基础作用。因此,当前的研究为理解深度学习中的 NGD 方法提供了新颖且统一的理论基础。
Natural Gradient Descent (NGD) helps to accelerate the convergence of gradient descent dynamics, but it requires approximations in large-scale deep neural networks because of its high computational cost. Empirical studies have confirmed that some NGD methods with approximate Fisher information converge sufficiently fast in practice. Nevertheless, it remains unclear from the theoretical perspective why and under what conditions such heuristic approximations work well. In this work, we reveal that, under specific conditions, NGD with approximate Fisher information achieves the same fast convergence to global minima as exact NGD. We consider deep neural networks in the infinite-width limit, and analyze the asymptotic training dynamics of NGD in function space via the neural tangent kernel. In the function space, the training dynamics with the approximate Fisher information are identical to those with the exact Fisher information, and they converge quickly. The fast convergence holds in layer-wise approximations; for instance, in block diagonal approximation where each block corresponds to a layer as well as in block tri-diagonal and K-FAC approximations. We also find that a unit-wise approximation achieves the same fast convergence under some assumptions. All of these different approximations have an isotropic gradient in the function space, and this plays a fundamental role in achieving the same convergence properties in training. Thus, the current study gives a novel and unified theoretical foundation with which to understand NGD methods in deep learning.