GradInit: Learning to Initialize Neural Networks for Stable and Efficient Training

GradInit: Learning to Initialize Neural Networks for Stable and Efficient Training
复制标题

DOI:
--
复制
发表时间:
2021-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Chen Zhu;Renkun Ni;Zheng Xu;Kezhi Kong;W. R. Huang;T. Goldstein
Chen Zhu;Renkun Ni;Zheng Xu;Kezhi Kong;W. R. Huang;T. Goldstein
中科院分区:
其他
文献类型:
--
作者:
Chen Zhu;Renkun Ni;Zheng Xu;Kezhi Kong;W. R. Huang;T. Goldstein

文献摘要

被引文献

相似文献

神经架构的创新促进了语言建模和计算机视觉的重大突破。不幸的是,如果网络参数没有正确初始化,新的架构通常会导致具有挑战性的超参数选择和训练不稳定性。已经提出了许多特定于体系结构的初始化方案,但这些方案并不总是可移植到新的体系结构。本文介绍了GradInit,一个自动化和架构无关的方法初始化神经网络。GradInit基于一个简单的启发式算法;调整每个网络层的范数,以便SGD或Adam的一个步骤与规定的超参数产生最小的可能损失值。这种调整是通过在每个参数块前面引入标量乘数变量,然后使用简单的数值方案优化这些变量来完成的。GradInit加速了许多卷积架构的收敛和测试性能,无论是否有跳过连接,甚至没有归一化层。它还提高了机器翻译的原始Transformer架构的稳定性,使其能够在广泛的学习率和动量系数下使用Adam或SGD进行训练,而无需学习率预热。代码可在https://github.com/zhuchen03/gradinit上获得。
Innovations in neural architectures have fostered significant breakthroughs in language modeling and computer vision. Unfortunately, novel architectures often result in challenging hyper-parameter choices and training instability if the network parameters are not properly initialized. A number of architecture-specific initialization schemes have been proposed, but these schemes are not always portable to new architectures. This paper presents GradInit, an automated and architecture agnostic method for initializing neural networks. GradInit is based on a simple heuristic; the norm of each network layer is adjusted so that a single step of SGD or Adam with prescribed hyperparameters results in the smallest possible loss value. This adjustment is done by introducing a scalar multiplier variable in front of each parameter block, and then optimizing these variables using a simple numerical scheme. GradInit accelerates the convergence and test performance of many convolutional architectures, both with or without skip connections, and even without normalization layers. It also improves the stability of the original Transformer architecture for machine translation, enabling training it without learning rate warmup using either Adam or SGD under a wide range of learning rates and momentum coefficients. Code is available at https://github.com/zhuchen03/gradinit.