On the Convergence of SGD with Biased Gradients
On the Convergence of SGD with Biased Gradients
复制标题
关于带有偏置梯度的 SGD 的收敛性
DOI:
--
复制
发表时间:
2020
期刊:
影响因子:
--
通讯作者:
Sebastian U. Stich
中科院分区:
文献类型:
--
作者:
Ahmad Ajalloeian;Sebastian U. Stich
We analyze the complexity of biased stochastic gradient methods (SGD), where individual updates are corrupted by deterministic, i.e. biased error terms. We derive convergence results for smooth (non-convex) functions and give improved rates under the Polyak-Łojasiewicz condition. We quantify how the magnitude of the bias impacts the attainable accuracy and the convergence rates (sometimes leading to divergence). Our framework covers many applications where either only biased gradient updates are available, or preferred, over unbiased ones for performance reasons. For instance, in the domain of distributed learning, biased gradient compression techniques such as top- k compression have been proposed as a tool to alleviate the communication bottleneck and in derivative-free optimization, only biased gradient estimators can be queried. We discuss a few guiding examples that show the broad appli-cability of our analysis.
影响因子:
2.7
作者:
Bin Hu;P. Seiler;Laurent Lessard
通讯作者:
Bin Hu;P. Seiler;Laurent Lessard
DOI:
--
发表时间:
2018-06
期刊:
ArXiv
影响因子:
--
作者:
Hongyi Wang;Scott Sievert;Shengchao Liu;Zachary B. Charles;Dimitris Papailiopoulos;S. Wright
通讯作者:
Hongyi Wang;Scott Sievert;Shengchao Liu;Zachary B. Charles;Dimitris Papailiopoulos;S. Wright