Bounding reward measures of Markov models using the Markov decision processes

Bounding reward measures of Markov models using the Markov decision processes
复制标题

使用马尔可夫决策过程限制马尔可夫模型的奖励措施

DOI:
10.1002/nla.792
复制
发表时间:
2011
影响因子:
4.3
通讯作者:
Buchholz
Buchholz
中科院分区:
数学3区
文献类型:
--
作者:
Buchholz

文献摘要

参考文献

被引文献

相似文献

对于马尔可夫报酬过程,在转移率和报酬的上下界已知的情况下,提出了一种新的期望报酬界的确定方法。基于以前的文件中,尖锐的界限已被定义的问题,但只有一个效率低下,不稳定的算法提出,本文提出了算法来计算的界限解释的问题作为一个马尔可夫决策过程。通过这种方式,可以采用众所周知的值和策略迭代算法来稳定且相当有效地计算奖励边界。不同的数值技术计算的奖励界限。版权所有© 2011约翰威利父子有限公司.
For a Markov reward process, where upper and lower bounds for the transition rates and rewards are known, a new approach to bound the expected reward is presented. Based on a previous paper where sharp bounds have been defined for the problem, but only an inefficient and unstable algorithm is proposed, this paper presents algorithms to compute the bounds by interpreting the problem as a Markov Decision Process. In this way, the well known value and policy iteration algorithms can be adopted to compute reward bounds in a stable and fairly efficient way. Different numerical techniques are presented for computing the reward bounds. Copyright © 2011 John Wiley & Sons, Ltd.
DOI: 10.1137/0913035
发表时间: 1992-03-01
期刊: SIAM JOURNAL ON SCIENTIFIC AND STATISTICAL COMPUTING
影响因子: --
作者:
VANDERVORST, HA
通讯作者: VANDERVORST, HA
DOI: 10.1145/179812.179848
发表时间: 1994-07
期刊: J. ACM
影响因子: --
作者:
John C.S. Lui;R. Muntz
通讯作者: John C.S. Lui;R. Muntz
DOI: 10.1145/1345263.1345360
发表时间: 2007-10
期刊: --
影响因子: --
作者:
J. Lambert;B. V. Houdt;C. Blondia
通讯作者: J. Lambert;B. V. Houdt;C. Blondia
DOI: --
发表时间: 1984
期刊: JACM
影响因子: --
作者:
P. Courtois;P. Semal
通讯作者: P. Semal
DOI: 10.1137/1.9780898718003
发表时间: 2003-05
期刊: --
影响因子: --
作者:
Y. Saad
通讯作者: Y. Saad