Step-Ahead Error Feedback for Distributed Training with Compressed Gradient

Step-Ahead Error Feedback for Distributed Training with Compressed Gradient
复制标题

DOI:
10.1609/aaai.v35i12.17254
复制
发表时间:
2020-08
期刊:
--
影响因子:
--
通讯作者:
An Xu;Zhouyuan Huo;Heng Huang
An Xu;Zhouyuan Huo;Heng Huang
中科院分区:
其他
文献类型:
--
作者:
An Xu;Zhouyuan Huo;Heng Huang

文献摘要

相似文献

尽管分布式机器学习方法可以加速大型深度神经网络的训练,但通信成本已成为制约性能的不可忽视的瓶颈。为了应对这一挑战,设计了基于梯度压缩的通信高效分布式学习方法来降低通信成本,并且最近结合了局部错误反馈来补偿相应的性能损失。然而,在本文中,我们将证明集中式分布式训练中的局部误差反馈会引发新的“梯度不匹配”问题,并且与全精度训练相比,可能会导致性能下降。为了解决这个关键问题,我们提出了两种新颖的技术,1)超前和2)误差平均,并进行了严格的理论分析。我们的理论和实证结果都表明我们的新方法可以处理“梯度不匹配”问题。实验结果表明,与训练时期的全精度训练和局部误差反馈相比,使用常见的梯度压缩方案,我们甚至可以训练得更快,并且没有性能损失。
Although the distributed machine learning methods can speed up the training of large deep neural networks, the communication cost has become the non-negligible bottleneck to constrain the performance. To address this challenge, the gradient compression based communication-efficient distributed learning methods were designed to reduce the communication cost, and more recently the local error feedback was incorporated to compensate for the corresponding performance loss. However, in this paper, we will show that a new "gradient mismatch" problem is raised by the local error feedback in centralized distributed training and can lead to degraded performance compared with full-precision training. To solve this critical problem, we propose two novel techniques, 1) step ahead and 2) error averaging, with rigorous theoretical analysis. Both our theoretical and empirical results show that our new methods can handle the "gradient mismatch" problem. The experimental results show that we can even train faster with common gradient compression schemes than both the full-precision training and local error feedback regarding the training epochs and without performance loss.