How does momentum benefit deep neural networks architecture design? A few case studies

How does momentum benefit deep neural networks architecture design? A few case studies
复制标题

DOI:
10.1007/s40687-022-00352-0
复制
发表时间:
2021-10
影响因子:
1.2
通讯作者:
Bao Wang;Hedi Xia;T. Nguyen;S. Osher
Bao Wang;Hedi Xia;T. Nguyen;S. Osher
中科院分区:
数学3区
文献类型:
--
作者:
Bao Wang;Hedi Xia;T. Nguyen;S. Osher

文献摘要

相似文献

我们提出并回顾了一种通过动量改进神经网络架构设计的算法和理论框架。作为案例研究,我们考虑动量如何改进循环神经网络 (RNN)、神经常微分方程 (ODE) 和 Transformer 的架构设计。我们证明,将动量集成到神经网络架构中具有几个显着的理论和经验优势,包括(1)将动量集成到 RNN 和神经 ODE 中可以克服训练 RNN 和神经 ODE 中的梯度消失问题,从而实现有效的学习长期依赖性; (2)神经常微分方程中的动量可以降低常微分方程动力学的刚度,从而显着提高训练和测试中的计算效率; (3)动量可以提高变压器的效率和精度。
We present and review an algorithmic and theoretical framework for improving neural network architecture design via momentum. As case studies, we consider how momentum can improve the architecture design for recurrent neural networks (RNNs), neural ordinary differential equations (ODEs), and transformers. We show that integrating momentum into neural network architectures has several remarkable theoretical and empirical benefits, including (1) integrating momentum into RNNs and neural ODEs can overcome the vanishing gradient issues in training RNNs and neural ODEs, resulting in effective learning long-term dependencies; (2) momentum in neural ODEs can reduce the stiffness of the ODE dynamics, which significantly enhances the computational efficiency in training and testing; (3) momentum can improve the efficiency and accuracy of transformers.