Arbitrarily modulated Markov decision processes
Arbitrarily modulated Markov decision processes
复制标题
任意调制的马尔可夫决策过程
DOI:
10.1109/cdc.2009.5400662
复制
发表时间:
2009
期刊:
影响因子:
--
通讯作者:
Shie Mannor
中科院分区:
文献类型:
--
作者:
Jia Yuan Yu;Shie Mannor
We consider decision-making problems in Markov decision processes where both the rewards and the transition probabilities vary in an arbitrary (e.g., nonstationary) fashion. We propose an online Q-learning style algorithm and give a guarantee on its performance evaluated in retrospect against alternative policies. Unlike previous works, the guarantee depends critically on the variability of the uncertainty in the transition probabilities, but holds regardless of arbitrary changes in rewards and transition probabilities over time. Besides its intrinsic computational efficiency, this approach requires neither prior knowledge nor estimation of the transition probabilities.