Technical Note - Markov Decision Processes with State-Information Lag

Technical Note - Markov Decision Processes with State-Information Lag
复制标题

技术说明 - 具有状态信息滞后的马尔可夫决策过程

DOI:
--
复制
发表时间:
1972
影响因子:
2.7
通讯作者:
C. Leondes
C. Leondes
中科院分区:
管理学4区
文献类型:
--
作者:
D. M. Brooks;C. Leondes

文献摘要

被引文献

相似文献

马尔可夫决策过程公式化提供了一种方法,用于在状态变化是马尔可夫的过程中选择最优策略,但假设在过程的每个阶段都有关于过程状态的完美信息。在实际过程状态的可用观测提供不完美的状态信息的情况下,马尔可夫决策过程方法仅在观测状态以马尔可夫方式变化时才适用。虽然这在一般情况下不是真的,但它确实适用于重要的特殊情况,即关于物理状态的信息在一个过渡或阶段的延迟之后变得可用。这个信息滞后过程可以被分析为马尔可夫决策过程。从完全信息过程的增益或单位时间的预期收益的退化提供了改进信息系统的潜在价值的度量。
The Markov-decision-process formulation provides a method for selecting the optimal policy in a process where changes of state are Markovian, but assumes perfect information as to process state at each stage of the process. Where the available observations of the actual process state provide imperfect state information, the Markov-decision-process approach is applicable only if the observed state changes in a Markovian fashion. Although this is not true in the general case, it does apply in the important special case where information about the physical state becomes available after a delay of one transition or stage. This information-lag process can be analyzed as a Markov decision process. The degradation in gain, or expected return per unit time, from that of the perfect-information process provides a measure of the potential value of improving the information system.