On the Empirical State-Action Frequencies in Markov Decision Processes Under General Policies
On the Empirical State-Action Frequencies in Markov Decision Processes Under General Policies
复制标题
一般政策下马尔可夫决策过程中的经验状态动作频率
DOI:
10.1287/moor.1050.0148
复制
发表时间:
2005
期刊:
影响因子:
--
通讯作者:
J. Tsitsiklis
中科院分区:
文献类型:
--
作者:
Shie Mannor;J. Tsitsiklis
We consider the empirical state-action frequencies and the empirical reward in weakly communicating finite-state Markov decision processes under general policies. We define a certain polytope and establish that every element of this polytope is the limit of the empirical frequency vector, under some policy, in a strong sense. Furthermore, we show that the probability of exceeding a given distance between the empirical frequency vector and the polytope decays exponentially with time under every policy. We provide similar results for vector-valued empirical rewards.