Slow and Stale Gradients Can Win the Race

Slow and Stale Gradients Can Win the Race
复制标题

DOI:
10.1109/jsait.2021.3103770
复制
发表时间:
2018-03
期刊:
IEEE Journal on Selected Areas in Information Theory
影响因子:
--
通讯作者:
Sanghamitra Dutta;Gauri Joshi;Soumyadip Ghosh;Parijat Dube;P. Nagpurkar
Sanghamitra Dutta;Gauri Joshi;Soumyadip Ghosh;Parijat Dube;P. Nagpurkar
中科院分区:
其他
文献类型:
--
作者:
Sanghamitra Dutta;Gauri Joshi;Soumyadip Ghosh;Parijat Dube;P. Nagpurkar

文献摘要

被引文献

相似文献

分布式随机梯度下降(SGD)在以同步方式运行时,在等待最慢的工作者(落后者)时会在运行时遇到延迟。异步方法可以缓解掉队现象,但会导致梯度失效,从而对收敛误差产生不利影响。在这项工作中,我们通过分析训练模型中的误差和实际训练时间(挂钟时间)之间的权衡,提出了一个新的理论特征来描述异步方法提供的加速比。我们工作中的主要新奇之处在于,我们的运行时分析考虑了随机掉队延迟,这有助于我们设计和比较在掉队和过时之间取得平衡的分布式SGD算法。在没有有界或指数延迟假设的情况下,我们还给出了一种新的误差收敛分析。最后,基于误差-运行时间折衷的理论描述,我们提出了一种在分布式SGD中逐步改变同步性的方法,并在CIFAR10数据集上展示了该方法的性能。
Distributed Stochastic Gradient Descent (SGD) when run in a synchronous manner, suffers from delays in runtime as it waits for the slowest workers (stragglers). Asynchronous methods can alleviate stragglers, but cause gradient staleness that can adversely affect the convergence error. In this work, we present a novel theoretical characterization of the speedup offered by asynchronous methods by analyzing the trade-off between the error in the trained model and the actual training runtime (wallclock time). The main novelty in our work is that our runtime analysis considers random straggling delays, which helps us design and compare distributed SGD algorithms that strike a balance between straggling and staleness. We also provide a new error convergence analysis of asynchronous SGD variants without bounded or exponential delay assumptions. Finally, based on our theoretical characterization of the error-runtime trade-off, we propose a method of gradually varying synchronicity in distributed SGD and demonstrate its performance on the CIFAR10 dataset.