The Linear Program approach in multi-chain Markov Decision Processes revisited

The Linear Program approach in multi-chain Markov Decision Processes revisited
复制标题

重新审视多链马尔可夫决策过程中的线性规划方法

DOI:
10.1007/bf01415752
复制
发表时间:
1995
期刊:
Zeitschrift für Operations Research
影响因子:
--
通讯作者:
F. Spieksma
F. Spieksma
中科院分区:
--
文献类型:
--
作者:
E. Altman;F. Spieksma

文献摘要

被引文献

相似文献

线性编程是解决马尔可夫决策过程(MDP)的重要且有用的工具。它的推导依赖于动态编程方法,该方法也用于解决MDP。但是,对于具有多个约束的马尔可夫决策过程,唯一可用的方法基于线性程序。本文的目的是研究与多链MDP有关的此类线性程序的某些方面。我们首先对文献中可用的线性程序中出现的决策变量进行随机解释。然后,我们显示了多限制的马尔可夫决策过程,可以从[9]中建议的线性程序从控制问题的等效不受约束的Lagrange公式获得。这显示了线性程序方法与Lagrange方法之间的连接,该方法以前仅用于单个约束[3,14,15]。
Linear Programming is known to be an important and useful tool for solving Markov Decision Processes (MDP). Its derivation relies on the Dynamic Programming approach, which also serves to solve MDP. However, for Markov Decision Processes with several constraints the only available methods are based on Linear Programs. The aim of this paper is to investigate some aspects of such Linear Programs, related to multi-chain MDPs. We first present a stochastic interpretation of the decision variables that appear in the Linear Programs available in the literature. We then show for the multi-constrained Markov Decision Process that the Linear Program suggested in [9] can be obtained from an equivalent unconstrained Lagrange formulation of the control problem. This shows the connection between the Linear Program approach and the Lagrange approach, that was previously used only for the case of a single constraint [3, 14, 15].