Distributed Online Linear Quadratic Control for Linear Time-invariant Systems

Distributed Online Linear Quadratic Control for Linear Time-invariant Systems
复制标题

DOI:
10.23919/acc50511.2021.9483391
复制
发表时间:
2020-09
期刊:
2021 American Control Conference (ACC)
影响因子:
--
通讯作者:
Ting-Jui Chang;Shahin Shahrampour
Ting-Jui Chang;Shahin Shahrampour
中科院分区:
其他
文献类型:
--
作者:
Ting-Jui Chang;Shahin Shahrampour

文献摘要

相似文献

经典线性二次 (LQ) 控制以线性时不变 (LTI) 系统为中心,其中控制状态对引入了具有时不变参数的二次成本。在线优化和控制领域的最新进展为研究 LQ 问题提供了新颖的工具,这些工具对时变成本参数具有鲁棒性。受这一系列研究的启发,我们研究了相同 LTI 系统的分布式在线 LQ 问题。考虑一个多代理网络,其中每个代理都被建模为 LTI 系统。 LTI 系统与解耦的、随时间变化的二次成本相关联,这些成本按顺序显示。网络的目标是使所有代理的控制序列与事后最好的中心化策略的控制序列竞争,这是由遗憾的概念所捕获的。我们开发了在线 LQ 算法的分布式变体,该算法运行分布式在线梯度下降,并投影到半定规划(SDP)以生成控制器。我们建立了一个后悔界限缩放作为有限时间范围的平方根,这意味着代理随着时间的增长达成共识。我们进一步提供数值实验来验证我们的理论结果。
Classical linear quadratic (LQ) control centers around linear time-invariant (LTI) systems, where the control-state pairs introduce a quadratic cost with time-invariant parameters. Recent advancement in online optimization and control has provided novel tools to study LQ problems that are robust to time-varying cost parameters. Inspired by this line of research, we study the distributed online LQ problem for identical LTI systems. Consider a multi-agent network where each agent is modeled as an LTI system. The LTI systems are associated with decoupled, time-varying quadratic costs that are revealed sequentially. The goal of the network is to make the control sequence of all agents competitive to that of the best centralized policy in hindsight, captured by the notion of regret. We develop a distributed variant of the online LQ algorithm, which runs distributed online gradient descent with a projection to a semi-definite programming (SDP) to generate controllers. We establish a regret bound scaling as the square root of the finite time-horizon, implying that agents reach consensus as time grows. We further provide numerical experiments verifying our theoretical result.