Bounding reward measures of Markov models using the Markov decision processes
Bounding reward measures of Markov models using the Markov decision processes
复制标题
使用马尔可夫决策过程限制马尔可夫模型的奖励措施
DOI:
10.1002/nla.792
复制
发表时间:
2011
影响因子:
4.3
通讯作者:
Buchholz
中科院分区:
文献类型:
--
作者:
Buchholz
For a Markov reward process, where upper and lower bounds for the transition rates and rewards are known, a new approach to bound the expected reward is presented. Based on a previous paper where sharp bounds have been defined for the problem, but only an inefficient and unstable algorithm is proposed, this paper presents algorithms to compute the bounds by interpreting the problem as a Markov Decision Process. In this way, the well known value and policy iteration algorithms can be adopted to compute reward bounds in a stable and fairly efficient way. Different numerical techniques are presented for computing the reward bounds. Copyright © 2011 John Wiley & Sons, Ltd.
登录
查看更多内容
DOI:
10.1137/0913035
发表时间:
1992-03-01
期刊:
SIAM JOURNAL ON SCIENTIFIC AND STATISTICAL COMPUTING
影响因子:
--
作者:
VANDERVORST, HA
通讯作者:
VANDERVORST, HA
DOI:
10.1145/179812.179848
发表时间:
1994-07
期刊:
J. ACM
影响因子:
--
作者:
John C.S. Lui;R. Muntz
通讯作者:
John C.S. Lui;R. Muntz
DOI:
10.1145/1345263.1345360
发表时间:
2007-10
期刊:
--
影响因子:
--
作者:
J. Lambert;B. V. Houdt;C. Blondia
通讯作者:
J. Lambert;B. V. Houdt;C. Blondia
DOI:
--
发表时间:
1984
期刊:
JACM
影响因子:
--
作者:
P. Courtois;P. Semal
通讯作者:
P. Semal
DOI:
10.1137/1.9780898718003
发表时间:
2003-05
期刊:
--
影响因子:
--
作者:
Y. Saad
通讯作者:
Y. Saad