On the Empirical State-Action Frequencies in Markov Decision Processes Under General Policies

On the Empirical State-Action Frequencies in Markov Decision Processes Under General Policies
复制标题

一般政策下马尔可夫决策过程中的经验状态动作频率

DOI:
10.1287/moor.1050.0148
复制
发表时间:
2005
期刊:
Math. Oper. Res.
影响因子:
--
通讯作者:
J. Tsitsiklis
J. Tsitsiklis
中科院分区:
--
文献类型:
--
作者:
Shie Mannor;J. Tsitsiklis

文献摘要

被引文献

相似文献

研究了一般策略下弱通信有限状态马尔可夫决策过程的经验状态-动作频率和经验报酬。我们定义了一个特定的多面体,并建立了这个多面体的每个元素是经验频率向量的极限,在某种政策下,在很强的意义上。此外,我们表明,超过一个给定的经验频率向量和多面体之间的距离的概率随时间呈指数衰减在每一个政策。我们提供了类似的结果向量值的经验奖励。
We consider the empirical state-action frequencies and the empirical reward in weakly communicating finite-state Markov decision processes under general policies. We define a certain polytope and establish that every element of this polytope is the limit of the empirical frequency vector, under some policy, in a strong sense. Furthermore, we show that the probability of exceeding a given distance between the empirical frequency vector and the polytope decays exponentially with time under every policy. We provide similar results for vector-valued empirical rewards.