Optimal policy for minimizing risk models in Markov decision processes

Optimal policy for minimizing risk models in Markov decision processes
复制标题

DOI:
10.1016/s0022-247x(02)00097-5
复制
发表时间:
2002-07
影响因子:
1.3
通讯作者:
Yoshio Ohtsubo;K. Toyonaga
Yoshio Ohtsubo;K. Toyonaga
中科院分区:
数学3区
文献类型:
--
作者:
Yoshio Ohtsubo;K. Toyonaga

文献摘要

被引文献

相似文献

研究了状态空间可数且报酬有界的折扣马氏决策过程的最小化风险问题。刻画了有限和无限水平情形下的最优值,给出了无限水平情形下最优策略存在的两个充分条件。这些条件与《白皮书》(1993)中的引理3密切相关,但正如吴和林(1999)所指出的那样,引理3是不正确的。我们得到了引理为真的条件,在此条件下,我们证明了存在最优策略。在另一种情况下,我们证明了最优值是某一最优性方程的唯一解,并且在瞬变集上存在最优策略。
We consider the minimizing risk problems in discounted Markov decisions processes with countable state space and bounded general rewards. We characterize optimal values for finite and infinite horizon cases and give two sufficient conditions for the existence of an optimal policy in an infinite horizon case. These conditions are closely connected with Lemma 3 in White (1993), which is not correct as Wu and Lin (1999) point out. We obtain a condition for the lemma to be true, under which we show that there is an optimal policy. Under another condition we show that an optimal value is a unique solution to some optimality equation and there is an optimal policy on a transient set.