Control of unknown linear systems with Thompson sampling

Control of unknown linear systems with Thompson sampling
复制标题

用汤普森采样控制未知线性系统

DOI:
10.1109/allerton.2017.8262873
复制
发表时间:
2017
期刊:
2017 55th Annual Allerton Conference on Communication, Control, and Computing (Allerton)
影响因子:
--
通讯作者:
Rahul Jain
Rahul Jain
中科院分区:
--
文献类型:
--
作者:
Ouyang Yi;Mukul Gagrani;Rahul Jain

文献摘要

被引文献

相似文献

针对系统参数未知的线性二次控制问题,提出了一种基于汤普森采样的学习算法。该算法被称为动态插曲汤普森采样(TSDE),其中两个停止准则决定了汤普森采样中动态插曲的长度。第一个停止标准控制插曲长度的增长速度。当样本协方差矩阵的行列式小于前一个值的一半时,触发第二个停止准则。我们证明了在先验分布的某些条件下,TSDE累积到时间T的期望(贝叶斯)遗憾以Õ(√T)为界。这里Õ(.)隐藏常数和对数因子。这是LQ控制中学习的期望后悔的第一个Õ(√T)边界。通过数值仿真验证了该方法的性能。
We propose a Thompson sampling based learning algorithm for the Linear Quadratic (LQ) control problem with unknown system parameters. The algorithm is called Thompson sampling with dynamic episodes (TSDE) where two stopping criteria determine the lengths of the dynamic episodes in Thompson sampling. The first stopping criterion controls the growth rate of episode length. The second stopping criterion is triggered when the determinant of the sample covariance matrix is less than half of the previous value. We show under some conditions on the prior distribution that the expected (Bayesian) regret of TSDE accumulated up to time T is bounded by Õ(√T). Here Õ (.) hides constants and logarithmic factors. This is the first Õ (√T) bound on expected regret of learning in LQ control. Numerical simulations are provided to illustrate the performance of TSDE.