Adaptive Communication Strategies to Achieve the Best Error-Runtime Trade-off in Local-Update SGD

Adaptive Communication Strategies to Achieve the Best Error-Runtime Trade-off in Local-Update SGD
复制标题

DOI:
--
复制
发表时间:
2018-10
期刊:
ArXiv
影响因子:
--
通讯作者:
Jianyu Wang;Gauri Joshi
Jianyu Wang;Gauri Joshi
中科院分区:
其他
文献类型:
--
作者:
Jianyu Wang;Gauri Joshi

文献摘要

被引文献

相似文献

大规模机器学习训练,特别是分布式随机梯度下降,需要对系统固有的可变性(如节点离散和随机通信延迟)具有鲁棒性。这项工作考虑了一个分布式训练框架,其中每个工作节点允许执行局部模型更新,并定期对结果模型进行平均。我们分析了相对于时钟时间(而不是迭代次数)的错误收敛的真实速度,并分析了它是如何受到平均频率的影响的。主要贡献是AdaComm的设计,这是一种自适应通信策略,从不频繁的平均开始以节省通信延迟并提高收敛速度,然后增加通信频率以实现低错误层。训练深度神经网络的严格实验表明,AdaComm可以比完全同步SGD节省3倍的时间,并且仍然达到相同的最终训练损失。
Large-scale machine learning training, in particular distributed stochastic gradient descent, needs to be robust to inherent system variability such as node straggling and random communication delays. This work considers a distributed training framework where each worker node is allowed to perform local model updates and the resulting models are averaged periodically. We analyze the true speed of error convergence with respect to wall-clock time (instead of the number of iterations), and analyze how it is affected by the frequency of averaging. The main contribution is the design of AdaComm, an adaptive communication strategy that starts with infrequent averaging to save communication delay and improve convergence speed, and then increases the communication frequency in order to achieve a low error floor. Rigorous experiments on training deep neural networks show that AdaComm can take $3 \times$ less time than fully synchronous SGD, and still reach the same final training loss.