Trading performance for stability in Markov decision processes

Trading performance for stability in Markov decision processes
复制标题

DOI:
10.1016/j.jcss.2016.09.009
复制
发表时间:
2017-03-01
影响因子:
1.1
通讯作者:
Kucera, Antonin
Kucera, Antonin
中科院分区:
计算机科学3区
文献类型:
--
作者:
Brazdil, Tomas;Chatterjee, Krishnendu;Kucera, Antonin

文献摘要

被引文献

相似文献

我们研究有限状态马尔可夫决策过程的控制器综合问题,其目标是优化预期的平均回报性能和稳定性(也称为可变性在文献中)。我们认为,用平均收益的统计方差来表达稳定性的基本概念有时是不够的,并提出了一个替代的定义。我们表明,一个战略,确保预期的平均回报和方差低于给定的界限需要随机化和记忆,在上述两个定义。然后,我们表明,找到这样一个战略的问题可以表示为一组约束。(C)2016作者爱思唯尔公司出版
We study controller synthesis problems for finite-state Markov decision processes, where the objective is to optimize the expected mean-payoff performance and stability (also known as variability in the literature). We argue that the basic notion of expressing the stability using the statistical variance of the mean payoff is sometimes insufficient, and propose an alternative definition. We show that a strategy ensuring both the expected mean payoff and the variance below given bounds requires randomization and memory, under both the above definitions. We then show that the problem of finding such a strategy can be expressed as a set of constraints. (C) 2016 The Authors. Published by Elsevier Inc.