Sync-Switch: Hybrid Parameter Synchronization for Distributed Deep Learning

Sync-Switch: Hybrid Parameter Synchronization for Distributed Deep Learning
复制标题

DOI:
10.1109/icdcs51616.2021.00057
复制
发表时间:
2021-04
期刊:
2021 IEEE 41st International Conference on Distributed Computing Systems (ICDCS)
影响因子:
--
通讯作者:
Shijian Li;Oren Mangoubi;Lijie Xu;Tian Guo
Shijian Li;Oren Mangoubi;Lijie Xu;Tian Guo
中科院分区:
其他
文献类型:
--
作者:
Shijian Li;Oren Mangoubi;Lijie Xu;Tian Guo

文献摘要

被引文献

相似文献

随机梯度下降(SGD)已成为培训分布式簇中深神经网络的事实上的方式。确定训练吞吐量和模型精度的关键因素是参数同步协议的选择。例如,虽然散装同步平行(BSP)通常达到更好的融合精度,但相应的训练吞吐量可能会受到散乱者的负面影响。相反,异步平行(ASP)可以具有较高的吞吐量,但其收敛性和准确性可能会受到陈旧梯度的影响。为了提高同步协议的性能,最近的工作通常着重于设计新协议,并严重依赖难以调整的超参数。在本文中,我们设计了一种混合同步方法,可利用BSP和ASP的好处,即减少训练时间,同时保持融合的准确性。基于广泛的经验分析,我们设计了一系列自适应策略的集合,这些策略决定了如何以及何时在同步协议之间切换。我们的政策包括针对反复出现的工作的脱机政策和用于处理瞬态散乱者的在线工作。我们在张力流的基础上,在称为同步开关的原型系统中实施了拟议的策略,并通过流行的深度学习模型和数据集评估了培训性能。我们的实验表明,同步开关可以实现ASP级训练的速度,同时在与BSP相比时保持相似的融合精度。此外,同步开关的基于弹性的策略可以充分减轻瞬态散乱者的影响。
Stochastic Gradient Descent (SGD) has become the de facto way to train deep neural networks in distributed clusters. A critical factor in determining the training throughput and model accuracy is the choice of the parameter synchronization protocol. For example, while Bulk Synchronous Parallel (BSP) often achieves better converged accuracy, the corresponding training throughput can be negatively impacted by stragglers. In contrast, Asynchronous Parallel (ASP) can have higher throughput, but its convergence and accuracy can be impacted by stale gradients. To improve the performance of synchronization protocol, recent work often focuses on designing new protocols with a heavy reliance on hard-to-tune hyper-parameters. In this paper, we design a hybrid synchronization approach that exploits the benefits of both BSP and ASP, i.e., reducing training time while simultaneously maintaining the converged accuracy. Based on extensive empirical profiling, we devise a collection of adaptive policies that determine how and when to switch between synchronization protocols. Our policies include both offline ones that target recurring jobs and online ones for handling transient stragglers. We implement the proposed policies in a prototype system, called Sync-Switch, on top of TensorFlow, and evaluate the training performance with popular deep learning models and datasets. Our experiments show that Sync-Switch can achieve ASP level training speedup while maintaining similar converged accuracy when comparing to BSP. Moreover, Sync-Switch's elastic-based policy can adequately mitigate the impact from transient stragglers.