An Algorithm to Identify and Compute Average Optimal Policies in Multichain Markov Decision Processes
An Algorithm to Identify and Compute Average Optimal Policies in Multichain Markov Decision Processes
复制标题
一种识别和计算多链马尔可夫决策过程中平均最优策略的算法
DOI:
10.1287/moor.28.3.553.16388
复制
发表时间:
2003
期刊:
影响因子:
--
通讯作者:
A. Leizarowitz
中科院分区:
文献类型:
--
作者:
A. Leizarowitz
This paper concerns discrete-time, finite state multichain MDPs with compact action sets. The optimality criterion is long-run average cost. Simple examples illustrate that optimal stationary Markov policies do not always exist. We establish the existence of e-optimal policies that are stationary Markovian, and develop an algorithm that computes these approximate optimal policies. We establish a necessary and sufficient condition for the existence of an optimal policy that is stationary Markovian, and in case that such an optimal policy exists the algorithm computes it.