An Algorithm to Identify and Compute Average Optimal Policies in Multichain Markov Decision Processes

An Algorithm to Identify and Compute Average Optimal Policies in Multichain Markov Decision Processes
复制标题

一种识别和计算多链马尔可夫决策过程中平均最优策略的算法

DOI:
10.1287/moor.28.3.553.16388
复制
发表时间:
2003
期刊:
Math. Oper. Res.
影响因子:
--
通讯作者:
A. Leizarowitz
A. Leizarowitz
中科院分区:
--
文献类型:
--
作者:
A. Leizarowitz

文献摘要

被引文献

相似文献

本文研究了离散时间、有限状态的多链MDP,其作用集是紧的。最优准则是长期平均成本。简单的例子说明,最佳的平稳马尔可夫策略并不总是存在的。我们建立了稳定的马尔可夫的e-最优策略的存在,并开发了一个算法,计算这些近似的最优策略。我们建立了一个必要和充分的条件,存在一个最佳的政策,是平稳的马尔可夫,并在这种情况下,这样的最佳政策存在的算法计算它。
This paper concerns discrete-time, finite state multichain MDPs with compact action sets. The optimality criterion is long-run average cost. Simple examples illustrate that optimal stationary Markov policies do not always exist. We establish the existence of e-optimal policies that are stationary Markovian, and develop an algorithm that computes these approximate optimal policies. We establish a necessary and sufficient condition for the existence of an optimal policy that is stationary Markovian, and in case that such an optimal policy exists the algorithm computes it.