Regret Analysis of Distributed Online LQR Control for Unknown LTI Systems

Regret Analysis of Distributed Online LQR Control for Unknown LTI Systems
复制标题

DOI:
10.1109/tac.2023.3299551
复制
发表时间:
2021-05
影响因子:
6.8
通讯作者:
Ting-Jui Chang;Shahin Shahrampour
Ting-Jui Chang;Shahin Shahrampour
中科院分区:
计算机科学2区
文献类型:
--
作者:
Ting-Jui Chang;Shahin Shahrampour

文献摘要

被引文献

相似文献

在线优化最近为研究事先未知的时变成本函数的最优控制开辟了新的途径。在这一研究思路的启发下,我们研究了具有未知动态的线性定常系统的分布式在线线性二次型调节器问题。考虑一个多代理网络,其中每个代理被建模为一个LTI系统。该网络具有全局时变的二次型代价,这种代价可能会相反地演化,并且只能被每个代理顺序地部分观察到。网络的目标是1)估计未知的动态,2)计算与事后最好的集中式策略竞争的本地控制序列,这将使网络成本的总和随着时间的推移而最小化。这个问题被表述为遗憾最小化。我们提出了在线LQR算法的一个分布式变体,其中代理在探索阶段计算他们的系统估计。然后,每个代理对半定规划应用分布式在线梯度下降,该半定规划的可行集基于代理系统估计。我们证明了,在很高的概率下,我们提出的算法的遗憾界为$O(T^{2/3}\logT)$,这意味着随着时间的推移,所有代理都是一致的。并给出了仿真结果,验证了我们的理论保证。
Online optimization has recently opened avenues to study optimal control for time-varying cost functions that are unknown in advance. Inspired by this line of research, we study the distributed online linear quadratic regulator (LQR) problem for linear time-invariant (LTI) systems with unknown dynamics. Consider a multiagent network where each agent is modeled as an LTI system. The network has a global time-varying quadratic cost, which may evolve adversarially and is only partially observed by each agent sequentially. The goal of the network is to collectively 1) estimate the unknown dynamics and 2) compute local control sequences competitive to the best centralized policy in hindsight, which minimizes the sum of network costs over time. This problem is formulated as a regret minimization. We propose a distributed variant of the online LQR algorithm, where agents compute their system estimates during an exploration stage. Each agent then applies distributed online gradient descent on a semidefinite programming whose feasible set is based on the agent system estimate. We prove that with high probability, the regret bound of our proposed algorithm scales as $O(T^{2/3}\log T)$, implying the consensus of all agents over time. We also provide simulation results verifying our theoretical guarantee.