Training Thinner and Deeper Neural Networks: Jumpstart Regularization

Training Thinner and Deeper Neural Networks: Jumpstart Regularization
复制标题

DOI:
10.1007/978-3-031-08011-1_23
复制
发表时间:
2022-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Carles Roger Riera Molina;Camilo Rey;Thiago Serra;Eloi Puertas;O. Pujol
Carles Roger Riera Molina;Camilo Rey;Thiago Serra;Eloi Puertas;O. Pujol
中科院分区:
其他
文献类型:
--
作者:
Carles Roger Riera Molina;Camilo Rey;Thiago Serra;Eloi Puertas;O. Pujol

文献摘要

被引文献

相似文献

当神经网络具有多层时,其表现力更强。反过来,传统的训练方法只有在深度不会导致诸如梯度爆炸或消失之类的数值问题时才会成功,而当层足够宽时,这种情况发生的频率会降低。然而,增加宽度以获得更大的深度需要使用更重的计算资源,并导致模型过度参数化。这些后续问题已通过量化和剪枝等模型压缩方法部分解决,其中一些方法依赖于损失函数的基于归一化的正则化,使大多数参数的影响可以忽略不计。在这项工作中,我们建议使用正则化来防止神经元死亡或变得线性,我们将这种技术称为跳跃启动正则化。与传统训练相比,我们获得了更薄、更深、最重要的是参数效率更高的神经网络。
Neural networks are more expressive when they have multiple layers. In turn, conventional training methods are only successful if the depth does not lead to numerical issues such as exploding or vanishing gradients, which occur less frequently when the layers are sufficiently wide. However, increasing width to attain greater depth entails the use of heavier computational resources and leads to overparameterized models. These subsequent issues have been partially addressed by model compression methods such as quantization and pruning, some of which relying on normalization-based regularization of the loss function to make the effect of most parameters negligible. In this work, we propose instead to use regularization for preventing neurons from dying or becoming linear, a technique which we denote asjumpstart regularization. In comparison to conventional training, we obtain neural networks that are thinner, deeper, and—most importantly—more parameter-efficient.