Optimal decision procedures for finite markov chains. Part I: Examples

Optimal decision procedures for finite markov chains. Part I: Examples
复制标题

有限马尔可夫链的最优决策过程。

DOI:
10.2307/1426039
复制
发表时间:
1973
影响因子:
1.2
通讯作者:
J. Bather
J. Bather
中科院分区:
数学4区
文献类型:
--
作者:
J. Bather

文献摘要

被引文献

相似文献

对于有限状态空间的离散时间马尔可夫过程,根据任意时刻所处的状态,从给定的集合中选择转移概率来控制。给定每个选择的即时成本,需要在无限的未来中最小化预期成本,而不贴现。各种技术进行审查的情况下,当有一个有限的可能的转移矩阵,并给出一个例子来说明由逆向归纳得出的政策序列的不可预测的行为。进一步的例子表明,当存在无穷多个转移矩阵时,现有的方法可能失效。提出了一种新的方法,根据他们的可访问性从彼此的状态分类的想法的基础上。
A Markov process in discrete time with a finite state space is controlled by choosing the transition probabilities from a prescribed set depending on the state occupied at any time. Given the immediate cost for each choice, it is required to minimise the expected cost over an infinite future, without discounting. Various techniques are reviewed for the case when there is a finite set of possible transition matrices and an example is given to illustrate the unpredictable behaviour of policy sequences derived by backward induction. Further examples show that the existing methods may break down when there is an infinite family of transition matrices. A new approach is suggested, based on the idea of classifying the states according to their accessibility from one another.