Simulation-based optimization of Markov reward processes

Simulation-based optimization of Markov reward processes
复制标题

DOI:
10.1109/cdc.1998.757861
复制
发表时间:
1998-12
期刊:
Proceedings of the 37th IEEE Conference on Decision and Control (Cat. No.98CH36171)
影响因子:
--
通讯作者:
P. Marbach;J. Tsitsiklis
P. Marbach;J. Tsitsiklis
中科院分区:
其他
文献类型:
--
作者:
P. Marbach;J. Tsitsiklis

文献摘要

被引文献

相似文献

我们提出了一种基于仿真的算法来优化马尔可夫奖励过程中依赖于一组参数的平均奖励。作为一种特殊情况,该方法适用于马尔可夫决策过程,其中优化发生在一组参数化策略中。该算法涉及单采样路径的仿真,可以在线实现。给出了一个收敛结果(概率为1)。
We propose a simulation-based algorithm for optimizing the average reward in a Markov reward process that depends on a set of parameters. As a special case, the method applies to Markov decision processes where optimization takes place within a parametrized set of policies. The algorithm involves the simulation of a single sample path, and can be implemented online. A convergence result (with probability 1) is provided.