Arbitrarily modulated Markov decision processes

Arbitrarily modulated Markov decision processes
复制标题

任意调制的马尔可夫决策过程

DOI:
10.1109/cdc.2009.5400662
复制
发表时间:
2009
期刊:
Proceedings of the 48h IEEE Conference on Decision and Control (CDC) held jointly with 2009 28th Chinese Control Conference
影响因子:
--
通讯作者:
Shie Mannor
Shie Mannor
中科院分区:
--
文献类型:
--
作者:
Jia Yuan Yu;Shie Mannor

文献摘要

被引文献

相似文献

我们考虑马尔可夫决策过程中的决策问题,其中奖励和转移概率都在任意(例如,非静止)方式。我们提出了一个在线Q-学习风格的算法,并保证其性能评估,在回顾对替代政策。与以前的作品不同,保证关键取决于在转移概率的不确定性的变化,但持有不管任意变化的奖励和转移概率随着时间的推移。除了其固有的计算效率,这种方法既不需要先验知识,也不需要估计的转移概率。
We consider decision-making problems in Markov decision processes where both the rewards and the transition probabilities vary in an arbitrary (e.g., nonstationary) fashion. We propose an online Q-learning style algorithm and give a guarantee on its performance evaluated in retrospect against alternative policies. Unlike previous works, the guarantee depends critically on the variability of the uncertainty in the transition probabilities, but holds regardless of arbitrary changes in rewards and transition probabilities over time. Besides its intrinsic computational efficiency, this approach requires neither prior knowledge nor estimation of the transition probabilities.