Optimal decision procedures for finite Markov chains. Part II: Communicating systems

Optimal decision procedures for finite Markov chains. Part II: Communicating systems
复制标题

有限马尔可夫链的最优决策过程。

DOI:
--
复制
发表时间:
1973
影响因子:
1.2
通讯作者:
J. Bather
J. Bather
中科院分区:
数学4区
文献类型:
--
作者:
J. Bather

文献摘要

被引文献

相似文献

有限状态空间中离散时间的马尔可夫过程是通过根据当前状态从给定的凸分布族中选择转移概率来控制的。每个选择的即时成本都是规定的,并且要求在无限的未来中最小化平均预期成本。本文考虑了这个一般问题的一个特例,并提供了一般解决方案的基础。主要结果是,存在一个最优策略,如果系统的每个状态都可以通过选择一个合适的策略从任何其他状态以正概率到达。
A Markov process in discrete time with a finite state space is controlled by choosing the transition probabilities from a given convex family of distributions depending on the present state. The immediate cost is prescribed for each choice and it is required to minimise the average expected cost over an infinite future. The paper considers a special case of this general problem and provides the foundation for a general solution. The main result is that an optimal policy exists if each state of the system can be reached with positive probability from any other state by choosing a suitable policy.