Optimal decision procedures for finite markov chains. Part I: Examples
Optimal decision procedures for finite markov chains. Part I: Examples
复制标题
有限马尔可夫链的最优决策过程。
DOI:
10.2307/1426039
复制
发表时间:
1973
影响因子:
1.2
通讯作者:
J. Bather
中科院分区:
文献类型:
--
作者:
J. Bather
A Markov process in discrete time with a finite state space is controlled by choosing the transition probabilities from a prescribed set depending on the state occupied at any time. Given the immediate cost for each choice, it is required to minimise the expected cost over an infinite future, without discounting. Various techniques are reviewed for the case when there is a finite set of possible transition matrices and an example is given to illustrate the unpredictable behaviour of policy sequences derived by backward induction. Further examples show that the existing methods may break down when there is an infinite family of transition matrices. A new approach is suggested, based on the idea of classifying the states according to their accessibility from one another.