Beyond BatchNorm: Towards a Unified Understanding of Normalization in Deep Learning

Beyond BatchNorm: Towards a Unified Understanding of Normalization in Deep Learning
复制标题

DOI:
--
复制
发表时间:
2021-06
期刊:
--
影响因子:
--
通讯作者:
Ekdeep Singh Lubana;R. Dick;Hidenori Tanaka
Ekdeep Singh Lubana;R. Dick;Hidenori Tanaka
中科院分区:
其他
文献类型:
--
作者:
Ekdeep Singh Lubana;R. Dick;Hidenori Tanaka

文献摘要

相似文献

受 BatchNorm 的启发,深度学习中的标准化层出现了爆炸式增长。最近的工作已经确定了 BatchNorm 的许多有益特性来解释其成功。然而,考虑到对替代标准化层的追求,需要对这些属性进行概括,以便可以准确预测任何给定层的成功/失败。在这项工作中,我们朝着这个目标迈出了第一步,将随机初始化的深度神经网络 (DNN) 中 BatchNorm 的已知属性扩展到最近提出的几个归一化层。我们的主要发现如下:(i)与 BatchNorm 类似,基于激活的归一化层可以防止 ResNet 中激活的指数增长,但参数技术需要明确的补救措施; (ii) 使用 GroupNorm 可以确保信息丰富的前向传播,不同的样本被分配不同的激活,但增加组大小会导致不同样本的激活越来越难以区分,这解释了使用 LayerNorm 的模型收敛速度慢的原因; (iii) 小组规模导致较早层中的梯度范数较大,因此解释了实例归一化中的训练不稳定问题,并说明了 GroupNorm 中的速度与稳定性权衡。总的来说,我们的分析揭示了一套统一的机制,这些机制支撑着深度学习中规范化方法的成功,为我们提供了一个指南针来系统地探索 DNN 规范化层的巨大设计空间。
Inspired by BatchNorm, there has been an explosion of normalization layers in deep learning. Recent works have identified a multitude of beneficial properties in BatchNorm to explain its success. However, given the pursuit of alternative normalization layers, these properties need to be generalized so that any given layer's success/failure can be accurately predicted. In this work, we take a first step towards this goal by extending known properties of BatchNorm in randomly initialized deep neural networks (DNNs) to several recently proposed normalization layers. Our primary findings follow: (i) similar to BatchNorm, activations-based normalization layers can prevent exponential growth of activations in ResNets, but parametric techniques require explicit remedies; (ii) use of GroupNorm can ensure an informative forward propagation, with different samples being assigned dissimilar activations, but increasing group size results in increasingly indistinguishable activations for different samples, explaining slow convergence speed in models with LayerNorm; and (iii) small group sizes result in large gradient norm in earlier layers, hence explaining training instability issues in Instance Normalization and illustrating a speed-stability tradeoff in GroupNorm. Overall, our analysis reveals a unified set of mechanisms that underpin the success of normalization methods in deep learning, providing us with a compass to systematically explore the vast design space of DNN normalization layers.