Dynamic Regret Minimization for Control of Non-stationary Linear Dynamical Systems

Dynamic Regret Minimization for Control of Non-stationary Linear Dynamical Systems
复制标题

非平稳线性动力系统控制的动态遗憾最小化

DOI:
10.1145/3508029
复制
发表时间:
2022
期刊:
Proceedings of the ACM on Measurement and Analysis of Computing Systems
影响因子:
--
通讯作者:
Kolar, Mladen
Kolar, Mladen
中科院分区:
--
文献类型:
--
作者:
Luo, Yuwei;Gupta, Varun;Kolar, Mladen

文献摘要

参考文献

被引文献

相似文献

考虑了线性二次型调节器(LQR)系统在有限水平T上的控制问题,该系统具有固定的和已知的代价矩阵Q,R,但未知的和非平稳的动态A_t,B_t,动态矩阵序列可以是任意的,但具有一个全变差V_T,假设V_T是控制器未知的。在假设所有t个都有一个可镇定但可能是次优的控制器序列的情况下,我们提出了一个算法来实现O(V_T^2/5 T^3/5)的最优动态遗憾。在分段恒定动态的情况下,我们的算法获得了O(SQRTST)的最优遗憾,其中S是开关的数目。我们算法的核心是一种自适应的非平稳性检测策略,该策略建立在最近针对上下文多臂Bandit问题开发的方法的基础上。我们还认为,对于LQR问题,非自适应遗忘(例如,重新开始或使用具有静态窗口大小的滑动窗口学习)可能不是最优的遗憾,即使在窗口大小被优化的情况下也是如此。我们算法分析中的主要技术挑战是证明当被估计的参数是非平稳的时,普通最小二乘(OLS)估计器具有小的偏差。我们的分析还强调,导致遗憾的关键主题是,LQR问题本质上是一个具有线性反馈和局部二次成本的强盗问题。这个主题比LQR问题本身更普遍,因此我们相信我们的结果应该得到更广泛的应用。
We consider the problem of controlling a Linear Quadratic Regulator (LQR) system over a finite horizon T with fixed and known cost matrices Q,R, but unknown and non-stationary dynamics A_t, B_t. The sequence of dynamics matrices can be arbitrary, but with a total variation, V_T, assumed to be o(T) and unknown to the controller. Under the assumption that a sequence of stabilizing, but potentially sub-optimal controllers is available for all t, we present an algorithm that achieves the optimal dynamic regret of O(V_T^2/5 T^3/5 ). With piecewise constant dynamics, our algorithm achieves the optimal regret of O(sqrtST ) where S is the number of switches. The crux of our algorithm is an adaptive non-stationarity detection strategy, which builds on an approach recently developed for contextual Multi-armed Bandit problems. We also argue that non-adaptive forgetting (e.g., restarting or using sliding window learning with a static window size) may not be regret optimal for the LQR problem, even when the window size is optimally tuned with the knowledge of. The main technical challenge in the analysis of our algorithm is to prove that the ordinary least squares (OLS) estimator has a small bias when the parameter to be estimated is non-stationary. Our analysis also highlights that the key motif driving the regret is that the LQR problem is in spirit a bandit problem with linear feedback and locally quadratic cost. This motif is more universal than the LQR problem itself, and therefore we believe our results should find wider application.
DOI: 10.1561/1700000031
发表时间: 2015
期刊: The American Economic Review
影响因子: --
作者:
P. Naik
通讯作者: P. Naik
用于自适应控制和学习的输入扰动
DOI: 10.1016/j.automatica.2020.108950
发表时间: 2020
期刊: Automatica
影响因子: 6.4
作者:
Shirani Faradonbeh, Mohamad Kazem;Tewari, Ambuj;Michailidis, George
通讯作者: Michailidis, George
DOI: 10.1287/moor.1090.0397
发表时间: 2008-11
期刊: Math. Oper. Res.
影响因子: --
作者:
Jia Yuan Yu;Shie Mannor;N. Shimkin
通讯作者: Jia Yuan Yu;Shie Mannor;N. Shimkin
DOI: 10.1145/1553374.1553425
发表时间: 2009-06
期刊: --
影响因子: --
作者:
Elad Hazan;Seshadhri Comandur
通讯作者: Elad Hazan;Seshadhri Comandur
DOI: --
发表时间: 2019
期刊: Conference on Uncertainty in Artificial Intelligence
影响因子: --
作者:
Pratik Gajane;R. Ortner;P. Auer
通讯作者: P. Auer