A Stochastic Linearized Augmented Lagrangian Method for Decentralized Bilevel Optimization

A Stochastic Linearized Augmented Lagrangian Method for Decentralized Bilevel Optimization
复制标题

DOI:
--
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Songtao Lu;Siliang Zeng;Xiaodong Cui;M. Squillante;L. Horesh;Brian Kingsbury;Jia Liu;Mingyi Hong
Songtao Lu;Siliang Zeng;Xiaodong Cui;M. Squillante;L. Horesh;Brian Kingsbury;Jia Liu;Mingyi Hong
中科院分区:
其他
文献类型:
--
作者:
Songtao Lu;Siliang Zeng;Xiaodong Cui;M. Squillante;L. Horesh;Brian Kingsbury;Jia Liu;Mingyi Hong

文献摘要

相似文献

双层优化已被证明是一种用于描述多任务机器学习问题的有效框架,例如强化学习(RL)和元学习,其中决策变量耦合在最小化问题的两个水平上。在实践中,学习任务将位于不同的计算资源环境中,因此需要部署一个分散的培训框架来实现多智能体和多任务学习。本文提出了一种随机线性化增广拉格朗日方法(SLAM)来求解图上一般的非凸双层优化问题,其中上下优化变量都能达成一致。我们还证明了所提出的SLAM算法对这类问题的Karush-Kuhn-Tucker(KKT)点的理论收敛速度与经典的分布随机梯度下降算法对单层非凸极小化问题的收敛速度是相同的。在多智能体RL问题上的数值测试结果表明了SLAM算法相对于基准算法的优越性。
Bilevel optimization has been shown to be a powerful framework for formulating multi-task machine learning problems, e.g., reinforcement learning (RL) and meta-learning, where the decision variables are coupled in both levels of the minimization problems. In practice, the learning tasks would be located at different computing resource environments, and thus there is a need for deploying a decentralized training framework to implement multi-agent and multi-task learning. We develop a stochastic linearized augmented Lagrangian method (SLAM) for solving general nonconvex bilevel optimization problems over a graph, where both upper and lower optimization variables are able to achieve a consensus. We also establish that the theoretical convergence rate of the proposed SLAM to the Karush-Kuhn-Tucker (KKT) points of this class of problems is on the same order as the one achieved by the classical distributed stochastic gradient descent for only single-level nonconvex minimization problems. Numerical results tested on multi-agent RL problems showcase the superiority of SLAM compared with the benchmarks.