On the Convergence of SGD with Biased Gradients

On the Convergence of SGD with Biased Gradients
复制标题

关于带有偏置梯度的 SGD 的收敛性

DOI:
--
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Sebastian U. Stich
Sebastian U. Stich
中科院分区:
--
文献类型:
--
作者:
Ahmad Ajalloeian;Sebastian U. Stich

文献摘要

参考文献

被引文献

相似文献

我们分析了有偏随机梯度方法(SGD)的复杂性,其中个别更新被破坏的确定性,即有偏误差项。我们推导出光滑(非凸)函数的收敛性结果,并给出改进的速度下的Polyak-Jasojasiewicz条件。我们量化了偏差的大小如何影响可达到的精度和收敛速度(有时会导致发散)。我们的框架涵盖了许多应用程序,其中只有偏置梯度更新可用,或首选,无偏的性能原因。例如,在分布式学习领域,已经提出了诸如top-k压缩的有偏梯度压缩技术作为缓解通信瓶颈的工具,并且在无导数优化中,只能查询有偏梯度估计器。我们讨论了一些指导性的例子,显示了我们的分析的广泛适用性。
We analyze the complexity of biased stochastic gradient methods (SGD), where individual updates are corrupted by deterministic, i.e. biased error terms. We derive convergence results for smooth (non-convex) functions and give improved rates under the Polyak-Łojasiewicz condition. We quantify how the magnitude of the bias impacts the attainable accuracy and the convergence rates (sometimes leading to divergence). Our framework covers many applications where either only biased gradient updates are available, or preferred, over unbiased ones for performance reasons. For instance, in the domain of distributed learning, biased gradient compression techniques such as top- k compression have been proposed as a tool to alleviate the communication bottleneck and in derivative-free optimization, only biased gradient estimators can be queried. We discuss a few guiding examples that show the broad appli-cability of our analysis.
DOI: 10.1007/s10107-020-01486-1
发表时间: 2017-11
影响因子: 2.7
作者:
Bin Hu;P. Seiler;Laurent Lessard
通讯作者: Bin Hu;P. Seiler;Laurent Lessard
DOI: --
发表时间: 2018-06
期刊: ArXiv
影响因子: --
作者:
Hongyi Wang;Scott Sievert;Shengchao Liu;Zachary B. Charles;Dimitris Papailiopoulos;S. Wright
通讯作者: Hongyi Wang;Scott Sievert;Shengchao Liu;Zachary B. Charles;Dimitris Papailiopoulos;S. Wright