Notes on equivalent stationary policies in Markov decision processes with total rewards
Notes on equivalent stationary policies in Markov decision processes with total rewards
复制标题
关于具有总奖励的马尔可夫决策过程中的等效固定策略的注释
DOI:
10.1007/bf01194331
复制
发表时间:
1996
影响因子:
1.2
通讯作者:
I. Sonin
中科院分区:
文献类型:
--
作者:
E. Feinberg;I. Sonin
We construct examples of Markov Decision Processes for which, for a given initial state and for a given nonstationary transient policy, there is no equivalent (randomized) stationary policy, i.e. there is no stationary policy which occupation measure is equal to the occupation measure of a given policy. We also investigate the relation between the existence of equivalent stationary policies in special models and the existence of equivalent strategies in various classes of nonstationary policies in general models.