Minibatch vs Local SGD for Heterogeneous Distributed Learning

Minibatch vs Local SGD for Heterogeneous Distributed Learning
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Blake E. Woodworth;Kumar Kshitij Patel;N. Srebro
Blake E. Woodworth;Kumar Kshitij Patel;N. Srebro
中科院分区:
其他
文献类型:
--
作者:
Blake E. Woodworth;Kumar Kshitij Patel;N. Srebro

文献摘要

被引文献

相似文献

我们在异构分布式环境中分析局部随机梯度下降(又名并行或联邦随机梯度下降)和小批量随机梯度下降。在该环境中,每台机器都能获取针对不同的、特定于机器的凸目标的随机梯度估计;目标是针对平均目标进行优化;并且机器只能间歇性地通信。我们认为:(i)在这种环境下,小批量随机梯度下降(即使没有加速)优于所有现有的局部随机梯度下降分析;(ii)当异构性较高时,加速的小批量随机梯度下降是最优的;(iii)给出了局部随机梯度下降的第一个上界,该上界在非同质情况下优于小批量随机梯度下降。
We analyze Local SGD (aka parallel or federated SGD) and Minibatch SGD in the heterogeneous distributed setting, where each machine has access to stochastic gradient estimates for a different, machine-specific, convex objective; the goal is to optimize w.r.t. the average objective; and machines can only communicate intermittently. We argue that, (i) Minibatch SGD (even without acceleration) dominates all existing analysis of Local SGD in this setting, (ii) accelerated Minibatch SGD is optimal when the heterogeneity is high, and (iii) present the first upper bound for Local SGD that improves over Minibatch SGD in a non-homogeneous regime.