Notes on equivalent stationary policies in Markov decision processes with total rewards

Notes on equivalent stationary policies in Markov decision processes with total rewards
复制标题

关于具有总奖励的马尔可夫决策过程中的等效固定策略的注释

DOI:
10.1007/bf01194331
复制
发表时间:
1996
影响因子:
1.2
通讯作者:
I. Sonin
I. Sonin
中科院分区:
数学4区
文献类型:
--
作者:
E. Feinberg;I. Sonin

文献摘要

被引文献

相似文献

We construct examples of Markov Decision Processes for which, for a given initial state and for a given nonstationary transient policy, there is no equivalent (randomized) stationary policy, i.e. there is no stationary policy which occupation measure is equal to the occupation measure of a given policy. We also investigate the relation between the existence of equivalent stationary policies in special models and the existence of equivalent strategies in various classes of nonstationary policies in general models.