A General Framework for Decentralized Optimization With First-Order Methods

A General Framework for Decentralized Optimization With First-Order Methods
复制标题

一阶方法求解分散优化问题的通用框架

DOI:
10.1109/jproc.2020.3024266
复制
发表时间:
2020-09
影响因子:
20.6
通讯作者:
Ran Xin;Shi Pu;Angelia Nedi'c;U. Khan
Ran Xin;Shi Pu;Angelia Nedi'c;U. Khan
中科院分区:
计算机科学1区
文献类型:
--
作者:
Ran Xin;Shi Pu;Angelia Nedi'c;U. Khan

文献摘要

相似文献

分散优化,以最小化的功能,分布在网络上的节点的有限和,一直是控制和信号处理研究的一个重要领域,由于其自然的最优控制和信号估计问题的相关性。最近,复杂计算和大规模数据科学需求的出现导致了这一领域活动的复苏。在这篇文章中,我们讨论了分散的一阶梯度方法,这些方法在控制,信号处理和机器学习问题中取得了巨大的成功,由于它们的简单性,这些方法是许多复杂推理和训练任务的首选方法。特别是,我们提供了一个一般的框架,分散的一阶方法,适用于有向和无向的通信网络一样,并表明,现有的大部分工作的优化和共识可以明确地与这个框架。我们进一步扩展的讨论分散随机一阶方法,依赖于随机梯度在每个节点和描述如何本地方差减少计划,以前被证明有希望在集中的设置,能够提高分散的方法的性能时,结合什么是所谓的梯度跟踪。我们激励和展示了相应方法在分散环境中出现的机器学习和信号处理问题的有效性。
Decentralized optimization to minimize a finite sum of functions, distributed over a network of nodes, has been a significant area within control and signal-processing research due to its natural relevance to optimal control and signal estimation problems. More recently, the emergence of sophisticated computing and large-scale data science needs have led to a resurgence of activity in this area. In this article, we discuss decentralized first-order gradient methods, which have found tremendous success in control, signal processing, and machine learning problems, where such methods, due to their simplicity, serve as the first method of choice for many complex inference and training tasks. In particular, we provide a general framework of decentralized first-order methods that is applicable to directed and undirected communication networks alike and show that much of the existing work on optimization and consensus can be related explicitly to this framework. We further extend the discussion to decentralized stochastic first-order methods that rely on stochastic gradients at each node and describe how local variance reduction schemes, previously shown to have promise in the centralized settings, are able to improve the performance of decentralized methods when combined with what is known as gradient tracking. We motivate and demonstrate the effectiveness of the corresponding methods in the context of machine learning and signal-processing problems that arise in decentralized environments.