Control of unknown linear systems with Thompson sampling
Control of unknown linear systems with Thompson sampling
复制标题
用汤普森采样控制未知线性系统
DOI:
10.1109/allerton.2017.8262873
复制
发表时间:
2017
期刊:
影响因子:
--
通讯作者:
Rahul Jain
中科院分区:
文献类型:
--
作者:
Ouyang Yi;Mukul Gagrani;Rahul Jain
We propose a Thompson sampling based learning algorithm for the Linear Quadratic (LQ) control problem with unknown system parameters. The algorithm is called Thompson sampling with dynamic episodes (TSDE) where two stopping criteria determine the lengths of the dynamic episodes in Thompson sampling. The first stopping criterion controls the growth rate of episode length. The second stopping criterion is triggered when the determinant of the sample covariance matrix is less than half of the previous value. We show under some conditions on the prior distribution that the expected (Bayesian) regret of TSDE accumulated up to time T is bounded by Õ(√T). Here Õ (.) hides constants and logarithmic factors. This is the first Õ (√T) bound on expected regret of learning in LQ control. Numerical simulations are provided to illustrate the performance of TSDE.