Why are Adaptive Methods Good for Attention Models?

Why are Adaptive Methods Good for Attention Models?
复制标题

DOI:
--
复制
发表时间:
2020-10
期刊:
arXiv: Optimization and Control
影响因子:
--
通讯作者:
J. Zhang;Sai Praneeth Karimireddy;Andreas Veit;Seungyeon Kim;Sashank J. Reddi;Sanjiv Kumar;S. Sra
J. Zhang;Sai Praneeth Karimireddy;Andreas Veit;Seungyeon Kim;Sashank J. Reddi;Sanjiv Kumar;S. Sra
中科院分区:
其他
文献类型:
--
作者:
J. Zhang;Sai Praneeth Karimireddy;Andreas Veit;Seungyeon Kim;Sashank J. Reddi;Sanjiv Kumar;S. Sra

文献摘要

被引文献

相似文献

虽然随机梯度下降(SGD)仍然是深度学习中的主要算法,但已经观察到像CLIPPED SGD/ADAM这样的自适应方法在重要任务(如注意模型)上的性能优于SGD。与自适应方法相比,SGD在什么情况下表现得更差还不是很清楚。在本文中,我们提供了经验和理论证据,证明噪声在随机梯度中的重尾分布是SGD性能较差的原因之一。给出了重尾噪声下自适应梯度算法的第一个紧致收敛上、下界。此外,我们还演示了渐变剪裁如何在解决重尾渐变噪声中起到关键作用。随后,我们通过开发一种坐标裁剪算法(ACClip)展示了裁剪如何在实践中应用,并展示了它在BERT预训练和精调任务中的优越性能。
While stochastic gradient descent (SGD) is still the \emph{de facto} algorithm in deep learning, adaptive methods like Clipped SGD/Adam have been observed to outperform SGD across important tasks, such as attention models. The settings under which SGD performs poorly in comparison to adaptive methods are not well understood yet. In this paper, we provide empirical and theoretical evidence that a heavy-tailed distribution of the noise in stochastic gradients is one cause of SGD's poor performance. We provide the first tight upper and lower convergence bounds for adaptive gradient methods under heavy-tailed noise. Further, we demonstrate how gradient clipping plays a key role in addressing heavy-tailed gradient noise. Subsequently, we show how clipping can be applied in practice by developing an \emph{adaptive} coordinate-wise clipping algorithm (ACClip) and demonstrate its superior performance on BERT pretraining and finetuning tasks.