Optimization Theory for ReLU Neural Networks Trained with Normalization Layers

Optimization Theory for ReLU Neural Networks Trained with Normalization Layers
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
--
影响因子:
--
通讯作者:
Yonatan Dukler;Quanquan Gu;Guido Montúfar
Yonatan Dukler;Quanquan Gu;Guido Montúfar
中科院分区:
其他
文献类型:
--
作者:
Yonatan Dukler;Quanquan Gu;Guido Montúfar

文献摘要

相似文献

作者:杜克勒 (Dukler)、尤纳坦 (Yonatan);顾泉泉;吉多·蒙图法尔 |摘要:深度神经网络的成功部分归功于归一化层的使用。批量归一化、层归一化和权重归一化等归一化层在实践中无处不在,因为它们提高了泛化性能并显着加快了训练速度。尽管如此,当前的绝大多数深度学习理论和非凸优化文献都关注非归一化设置,其中所考虑的函数不表现出常见归一化神经网络的属性。在本文中,我们通过给出使用归一化层(即权重归一化)训练的 ReLU 激活的两层神经网络的第一个全局收敛结果来弥补这一差距。我们的分析表明,与非归一化神经网络相比,归一化层的引入如何改变优化环境,并且可以实现更快的收敛。
Author(s): Dukler, Yonatan; Gu, Quanquan; Montufar, Guido | Abstract: The success of deep neural networks is in part due to the use of normalization layers. Normalization layers like Batch Normalization, Layer Normalization and Weight Normalization are ubiquitous in practice, as they improve generalization performance and speed up training significantly. Nonetheless, the vast majority of current deep learning theory and non-convex optimization literature focuses on the un-normalized setting, where the functions under consideration do not exhibit the properties of commonly normalized neural networks. In this paper, we bridge this gap by giving the first global convergence result for two-layer neural networks with ReLU activations trained with a normalization layer, namely Weight Normalization. Our analysis shows how the introduction of normalization layers changes the optimization landscape and can enable faster convergence as compared with un-normalized neural networks.