Batch Normalization Preconditioning for Neural Network Training

Batch Normalization Preconditioning for Neural Network Training
复制标题

DOI:
--
复制
发表时间:
2021-08
期刊:
ArXiv
影响因子:
--
通讯作者:
Susanna Lange;Kyle E. Helfrich;Qiang Ye
Susanna Lange;Kyle E. Helfrich;Qiang Ye
中科院分区:
其他
文献类型:
--
作者:
Susanna Lange;Kyle E. Helfrich;Qiang Ye

文献摘要

相似文献

批量归一化(BN)是深度学习中一种流行且普遍存在的方法,已被证明可以减少训练时间并提高神经网络的泛化性能。尽管国阵取得了成功,但理论上并没有得到很好的理解。它不适合用于非常小的小批量或在线学习。在本文中,我们提出了一种新的方法称为批量归一化预处理(BNP)。BNP不是像BN那样通过批量归一化层显式地应用归一化,而是通过在训练期间直接调节参数梯度来应用归一化。这是为了改善损失函数的Hessian矩阵,从而在训练期间收敛。一个好处是BNP不受小批量大小的限制,并且可以在在线学习环境中工作。此外,它与BN的连接提供了关于BN如何改进训练以及BN如何应用于卷积神经网络等特殊架构的理论见解。作为理论基础,我们还提出了一种新的基于Hessian条件数的收敛理论,该理论适用于具有尺度不变性质的网络。
Batch normalization (BN) is a popular and ubiquitous method in deep learning that has been shown to decrease training time and improve generalization performance of neural networks. Despite its success, BN is not theoretically well understood. It is not suitable for use with very small mini-batch sizes or online learning. In this paper, we propose a new method called Batch Normalization Preconditioning (BNP). Instead of applying normalization explicitly through a batch normalization layer as is done in BN, BNP applies normalization by conditioning the parameter gradients directly during training. This is designed to improve the Hessian matrix of the loss function and hence convergence during training. One benefit is that BNP is not constrained on the mini-batch size and works in the online learning setting. Furthermore, its connection to BN provides theoretical insights on how BN improves training and how BN is applied to special architectures such as convolutional neural networks. For a theoretical foundation, we also present a novel Hessian condition number based convergence theory for a locally convex but not strong-convex loss, which is applicable to networks with a scale-invariant property.