Bayesian Optimal Control of Smoothly Parameterized Systems

Bayesian Optimal Control of Smoothly Parameterized Systems
复制标题

DOI:
--
复制
发表时间:
2015-07
期刊:
--
影响因子:
--
通讯作者:
Yasin Abbasi-Yadkori;Csaba Szepesvari
Yasin Abbasi-Yadkori;Csaba Szepesvari
中科院分区:
其他
文献类型:
--
作者:
Yasin Abbasi-Yadkori;Csaba Szepesvari

文献摘要

被引文献

相似文献

研究了一类光滑参数化马尔可夫决策问题的贝叶斯最优控制问题。我们提出了一个懒惰的版本,所谓的后验抽样方法,一种方法,可以追溯到汤普森和斯特伦斯,最近由Osband,Russo和货车罗伊研究。虽然Osband等人.推导出了一个界限上的(贝叶斯)遗憾的这种方法的未贴现的总成本情节,有限的状态和行动的问题,我们认为连续的,平均成本设置没有基数限制的状态或行动空间。虽然在情节设置中,在情节结束时切换到新政策是很自然的,但在持续的平均成本框架中,我们必须明确地以原则性的方式引入切换点,否则遗憾可能会线性增长。我们的懒惰方法基于监控未知参数的不确定性来引入这些切换点。为了开发一个合适的和易于计算的不确定性度量,我们引入了一个新的“平均局部光滑”的条件,这是在常见的例子中得到满足。在这种情况下,和一些额外的温和的条件下,我们得出率最优的边界上的遗憾,我们的算法。我们的一般方法使我们能够使用一个单一的算法和一个单一的分析范围广泛的问题,如有限的MDP或线性二次调节,都是顺利参数化的MDP的实例。通过一个仿真例子说明了该方法的有效性。
We study Bayesian optimal control of a general class of smoothly parameterized Markov decision problems (MDPs). We propose a lazy version of the so-called posterior sampling method, a method that goes back to Thompson and Strens, more recently studied by Osband, Russo and van Roy. While Osband et al. derived a bound on the (Bayesian) regret of this method for undiscounted total cost episodic, finite state and action problems, we consider the continuing, average cost setting with no cardinality restrictions on the state or action spaces. While in the episodic setting, it is natural to switch to a new policy at the episode-ends, in the continuing average cost framework we must introduce switching points explicitly and in a principled fashion, or the regret could grow linearly. Our lazy method introduces these switching points based on monitoring the uncertainty left about the unknown parameter. To develop a suitable and easy-to-compute uncertainty measure, we introduce a new "average local smoothness" condition, which is shown to be satisfied in common examples. Under this, and some additional mild conditions, we derive rate-optimal bounds on the regret of our algorithm. Our general approach allows us to use a single algorithm and a single analysis for a wide range of problems, such as finite MDPs or linear quadratic regulation, both being instances of smoothly parameterized MDPs. The effectiveness of our method is illustrated by means of a simulated example.