A relaxed technical assumption for posterior sampling-based reinforcement learning for control of unknown linear systems

A relaxed technical assumption for posterior sampling-based reinforcement learning for control of unknown linear systems
复制标题

用于控制未知线性系统的基于后验采样的强化学习的宽松技术假设

DOI:
--
复制
发表时间:
2021
期刊:
arXiv.org
影响因子:
--
通讯作者:
Ouyang Yi
Ouyang Yi
中科院分区:
--
文献类型:
--
作者:
Mukul Gagrani;Sagar Sudhakara;Aditya Mahajan;A. Nayyar;Ouyang Yi

文献摘要

被引文献

相似文献

我们重新研究了Thompson采样算法,以控制Ouyang等人最近提出的未知线性二次(LQ)系统。在闭环系统诱导范数的技术假设下,导出了算法的遗憾界。在本技术说明中,我们表明,通过对算法进行轻微修改(特别是,确保情节不会过早结束),可以用闭环系统谱半径方面的较温和的假设来取代对诱导范数的技术假设。改进后的算法具有相同的贝叶斯遗憾度(≈O(√T)),其中T是时间范围,而≈O(·)符号隐藏了T中的对数项。
—We revisit the Thompson sampling algorithm to control an unknown linear quadratic (LQ) system recently proposed by Ouyang et al. [1]. The regret bound of the algorithm was derived under a technical assumption on the induced norm of the closed loop system. In this technical note, we show that by making a minor modification in the algorithm (in particular, ensuring that an episode does not end too soon), this technical assumption on the induced norm can be replaced by a milder assumption in terms of the spectral radius of the closed loop system. The modified algorithm has the same Bayesian regret of ˜ O ( √ T ) , where T is the time-horizon and the ˜ O ( · ) notation hides logarithmic terms in T .