Pseudometrics for State Aggregation in Average Reward Markov Decision Processes

Pseudometrics for State Aggregation in Average Reward Markov Decision Processes
复制标题

平均奖励马尔可夫决策过程中状态聚合的伪计量学

DOI:
10.1007/978-3-540-75225-7_30
复制
发表时间:
2007
期刊:
2019 18th European Control Conference (ECC)
影响因子:
--
通讯作者:
R. Ortner
R. Ortner
中科院分区:
--
文献类型:
--
作者:
R. Ortner

文献摘要

被引文献

相似文献

我们考虑了如何用伪度量来描述平均奖励马尔可夫决策过程(mdp)中的状态相似性。通过引入与MDP结构相适应的适当度量概念,我们展示了如何将这些概念用于状态聚合。我们给出了处理聚合MDP而不是原始MDP可能造成的损失的上限,并将其与贴现奖励MDP所达到的上限进行了比较。
We consider how state similarity in average reward Markov decision processes (MDPs) may be described by pseudometrics. Introducing the notion of adequatepseudometrics which are well adapted to the structure of the MDP, we show how these may be used for state aggregation. Upper bounds on the loss that may be caused by working on the aggregated instead of the original MDP are given and compared to the bounds that have been achieved for discounted reward MDPs.