A relaxed technical assumption for posterior sampling-based reinforcement learning for control of unknown linear systems
A relaxed technical assumption for posterior sampling-based reinforcement learning for control of unknown linear systems
复制标题
用于控制未知线性系统的基于后验采样的强化学习的宽松技术假设
DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Ouyang Yi
中科院分区:
文献类型:
--
作者:
Mukul Gagrani;Sagar Sudhakara;Aditya Mahajan;A. Nayyar;Ouyang Yi
—We revisit the Thompson sampling algorithm to control an unknown linear quadratic (LQ) system recently proposed by Ouyang et al. [1]. The regret bound of the algorithm was derived under a technical assumption on the induced norm of the closed loop system. In this technical note, we show that by making a minor modification in the algorithm (in particular, ensuring that an episode does not end too soon), this technical assumption on the induced norm can be replaced by a milder assumption in terms of the spectral radius of the closed loop system. The modified algorithm has the same Bayesian regret of ˜ O ( √ T ) , where T is the time-horizon and the ˜ O ( · ) notation hides logarithmic terms in T .