Optimal threshold probability in undiscounted Markov decision processes with a target set

Optimal threshold probability in undiscounted Markov decision processes with a target set
复制标题

DOI:
10.1016/s0096-3003(03)00158-9
复制
发表时间:
2004-02
期刊:
Appl. Math. Comput.
影响因子:
--
通讯作者:
Yoshio Ohtsubo
Yoshio Ohtsubo
中科院分区:
其他
文献类型:
--
作者:
Yoshio Ohtsubo

文献摘要

被引文献

相似文献

我们考虑风险最小化问题的非折扣马尔可夫决策过程的目标集。我们制定的问题作为一个无限的地平线的情况下,经常性的类。我们证明了最优值函数是最优性方程的唯一解,并且存在平稳最优策略。给出了几种数值迭代方法和一种策略改进方法。
We consider risk minimizing problems in undiscounted Markov decisions processes with a target set. We formulate the problem as an infinite horizon case with a recurrent class. We show that an optimal value function is a unique solution to an optimality equation and there exists an stationary optimal policy. Also we give several value iteration methods and a policy improvement method.