Robust Modified Policy Iteration

Robust Modified Policy Iteration
复制标题

稳健的修改策略迭代

DOI:
--
复制
发表时间:
2013
影响因子:
2.1
通讯作者:
A. Schaefer
A. Schaefer
中科院分区:
计算机科学3区
文献类型:
--
作者:
David L. Kaufman;A. Schaefer

文献摘要

被引文献

相似文献

强大的动态编程可降低过渡可能性对马尔可夫决策问题的影响修改的策略迭代rmpi并演示其收敛性。内部问题在指定的公差内解决了。
Robust dynamic programming robust DP mitigates the effects of ambiguity in transition probabilities on the solutions of Markov decision problems. We consider the computation of robust DP solutions for discrete-stage, infinite-horizon, discounted problems with finite state and action spaces. We present robust modified policy iteration RMPI and demonstrate its convergence. RMPI encompasses both of the previously known algorithms, robust value iteration and robust policy iteration. In addition to proposing exact RMPI, in which the “inner problem” is solved precisely, we propose inexact RMPI, in which the inner problem is solved to within a specified tolerance. We also introduce new stopping criteria based on the span seminorm. Finally, we demonstrate through some numerical studies that RMPI can significantly reduce computation time.