Stochastic optimization of controlled partially observable Markov decision processes
Stochastic optimization of controlled partially observable Markov decision processes
复制标题
DOI:
10.1109/cdc.2000.912744
复制
发表时间:
2000-12
期刊:
影响因子:
--
通讯作者:
P. Bartlett;Jonathan Baxter
中科院分区:
文献类型:
--
作者:
P. Bartlett;Jonathan Baxter
We introduce an online algorithm for finding local maxima of the average reward in a partially observable Markov decision process (POMDP) controlled by a parameterized policy. Optimization is over the parameters of the policy. The algorithm's chief advantages are that it requires only a single sample path of the POMDP, it uses only one free parameter /spl beta//spl isin/(0, 1), which has a natural interpretation in terms of a bias-variance trade-off, and it requires no knowledge of the underlying state. In addition, the algorithm can be applied to infinite state, control and observation spaces. We prove almost-sure convergence of our algorithm, and show how the correct setting of /spl beta/ is related to the mixing time of the Markov chain induced by the POMDP.