Stochastic optimization of controlled partially observable Markov decision processes

Stochastic optimization of controlled partially observable Markov decision processes
复制标题

DOI:
10.1109/cdc.2000.912744
复制
发表时间:
2000-12
期刊:
Proceedings of the 39th IEEE Conference on Decision and Control (Cat. No.00CH37187)
影响因子:
--
通讯作者:
P. Bartlett;Jonathan Baxter
P. Bartlett;Jonathan Baxter
中科院分区:
其他
文献类型:
--
作者:
P. Bartlett;Jonathan Baxter

文献摘要

被引文献

相似文献

我们引入了一种在线算法,用于在由参数化策略控制的部分可观察马尔可夫决策过程(POMDP)中找到平均奖励的局部最大值。优化是针对策略参数的。该算法的主要优点是它只需要 POMDP 的单个样本路径,它只使用一个自由参数 /spl beta//spl isin/(0, 1),它在偏差-方差权衡方面具有自然的解释,并且不需要了解底层状态。此外,该算法还可应用于无限状态、控制和观察空间。我们证明了我们的算法几乎肯定的收敛性,并展示了 /spl beta/ 的正确设置如何与 POMDP 引起的马尔可夫链的混合时间相关。
We introduce an online algorithm for finding local maxima of the average reward in a partially observable Markov decision process (POMDP) controlled by a parameterized policy. Optimization is over the parameters of the policy. The algorithm's chief advantages are that it requires only a single sample path of the POMDP, it uses only one free parameter /spl beta//spl isin/(0, 1), which has a natural interpretation in terms of a bias-variance trade-off, and it requires no knowledge of the underlying state. In addition, the algorithm can be applied to infinite state, control and observation spaces. We prove almost-sure convergence of our algorithm, and show how the correct setting of /spl beta/ is related to the mixing time of the Markov chain induced by the POMDP.