Detached Error Feedback for Distributed SGD with Random Sparsification

Detached Error Feedback for Distributed SGD with Random Sparsification
复制标题

DOI:
--
复制
发表时间:
2020-04
期刊:
--
影响因子:
--
通讯作者:
An Xu;Heng Huang
An Xu;Heng Huang
中科院分区:
其他
文献类型:
--
作者:
An Xu;Heng Huang

文献摘要

相似文献

通信瓶颈一直是大规模分布式深度学习的关键问题。在这项工作中,我们研究了以随机分块稀疏化作为梯度压缩器的分布式 SGD,它与环全归约兼容且计算效率高,但导致性能较差。为了解决这个重要问题,我们从一个新颖的方面改进了通信效率的分布式SGD,即梯度方差和二阶矩之间的权衡。出于这个动机,我们提出了一种新的分离误差反馈(DEF)算法,该算法在非凸问题上表现出比误差反馈更好的收敛界限。我们还提出 DEF-A 在训练的早期阶段加速 DEF 的泛化,这比 DEF 表现出更好的泛化界限。此外,我们首次在高效通信的分布式 SGD 和迭代平均 SGD (SGD-IA) 之间建立了联系。广泛的深度学习实验表明所提出的方法在各种设置下都有显着的经验改进。
The communication bottleneck has been a critical problem in large-scale distributed deep learning. In this work, we study distributed SGD with random block-wise sparsification as the gradient compressor, which is ring-allreduce compatible and highly computation-efficient but leads to inferior performance. To tackle this important issue, we improve the communication-efficient distributed SGD from a novel aspect, that is, the trade-off between the variance and second moment of the gradient. With this motivation, we propose a new detached error feedback (DEF) algorithm, which shows better convergence bound than error feedback for non-convex problems. We also propose DEF-A to accelerate the generalization of DEF at the early stages of the training, which shows better generalization bounds than DEF. Furthermore, we establish the connection between communication-efficient distributed SGD and SGD with iterate averaging (SGD-IA) for the first time. Extensive deep learning experiments show significant empirical improvement of the proposed methods under various settings.