Distributed SGD Generalizes Well Under Asynchrony

Distributed SGD Generalizes Well Under Asynchrony
复制标题

DOI:
10.1109/allerton.2019.8919791
复制
发表时间:
2019-09
期刊:
2019 57th Annual Allerton Conference on Communication, Control, and Computing (Allerton)
影响因子:
--
通讯作者:
Jayanth Reddy Regatti;Gaurav Tendolkar;Yi Zhou;Abhishek K. Gupta;Yingbin Liang
Jayanth Reddy Regatti;Gaurav Tendolkar;Yi Zhou;Abhishek K. Gupta;Yingbin Liang
中科院分区:
其他
文献类型:
--
作者:
Jayanth Reddy Regatti;Gaurav Tendolkar;Yi Zhou;Abhishek K. Gupta;Yingbin Liang

文献摘要

相似文献

在大数据时代,异步分布式系统以其强大的可扩展性成为主流,全同步分布式系统的性能面临瓶颈。本文研究了随机梯度下降算法在分布式异步系统上的推广性能。该系统由计算随机梯度的多个工作机器组成,这些随机梯度被进一步发送到公共参数服务器并在公共参数服务器上聚合以更新变量,并且系统中的通信遭受可能的延迟。在算法稳定性框架下,我们证明了分布式异步SGD在训练优化中,在足够的数据样本下具有很好的推广性。特别是,我们的研究结果表明,降低学习率,因为我们允许更多的分布式系统中的冗余。这种自适应学习率策略提高了分布式算法的稳定性,降低了相应的泛化误差。然后,我们通过数值实验证实了我们的理论研究结果。
The performance of fully synchronized distributed systems has faced a bottleneck due to the big data trend, under which asynchronous distributed systems are becoming a major popularity due to their powerful scalability. In this paper, we study the generalization performance of stochastic gradient descent (SGD) on a distributed asynchronous system. The system consists of multiple worker machines that compute stochastic gradients which are further sent to and aggregated on a common parameter server to update the variables, and the communication in the system suffers from possible delays. Under the algorithm stability framework, we prove that distributed asynchronous SGD generalizes well given enough data samples in the training optimization. In particular, our results suggest to reduce the learning rate as we allow more asynchrony in the distributed system. Such adaptive learning rate strategy improves the stability of the distributed algorithm and reduces the corresponding generalization error. Then, we confirm our theoretical findings via numerical experiments.