Randomized and Relaxed Strategies in Continuous-Time Markov Decision Processes

Randomized and Relaxed Strategies in Continuous-Time Markov Decision Processes
复制标题

DOI:
10.1137/15m1014012
复制
发表时间:
2015-12
期刊:
SIAM J. Control. Optim.
影响因子:
--
通讯作者:
A. Piunovskiy
A. Piunovskiy
中科院分区:
其他
文献类型:
--
作者:
A. Piunovskiy

文献摘要

相似文献

本文的目标之一是描述一个广泛的控制策略,其中包括传统的放松策略,以及所谓的随机策略,出现在半马尔可夫决策过程的框架。如果目标是跳跃累积的总期望成本,则不失一般性,可以只考虑马尔可夫松弛策略。在一个简单的条件下,马尔可夫随机化策略也是充分的。通过实例说明了上述条件的重要性。最后,在没有任何条件的情况下,这类所谓的Poisson相关策略在最优化问题中也是充分的。所得结果不仅适用于贴现模型,也适用于长期平均成本模型。
One of the goals of this article is to describe a wide class of control strategies, which includes the traditional relaxed strategies, as well as the so called randomized strategies which appeared earlier only in the framework of semi-Markov decision processes. If the objective is the total expected cost up to the accumulation of jumps, then without loss of generality one can consider only Markov relaxed strategies. Under a simple condition, the Markov randomized strategies are also sufficient. An example shows that the mentioned condition is important. Finally, without any conditions, the class of so called Poisson-related strategies is also sufficient in the optimization problems. All the results are applicable to the discounted model, they may be useful also for the case of long-run average cost.