Estimation and optimal control for constrained Markov chains

Estimation and optimal control for constrained Markov chains
复制标题

约束马尔可夫链的估计和最优控制

DOI:
--
复制
发表时间:
1986
期刊:
IEEE Conference on Decision and Control
影响因子:
--
通讯作者:
A. Shwartz
A. Shwartz
中科院分区:
--
文献类型:
--
作者:
Dye;A. Makowski;A. Shwartz

文献摘要

被引文献

相似文献

许多工程系统的(最优)设计可充分地重新表述为马尔可夫决策过程,其中对系统性能的要求以约束条件的形式体现。在本文中,对受约束的马尔可夫决策过程的各种最优性结果进行了简要回顾;讨论了相应的实施问题,并表明这些问题导致了若干参数估计问题。在排队系统的背景下给出了这种受约束问题自然产生的简单情形,以说明该理论的各个要点。在每种情况下,都展示了最优策略的结构。
The (optimal) design of many engineering systems can be adequately recast as a Markov decision process, where requirements on system performance are captured in the form of constraints. In this paper, various optimality results for constrained Markov decision processes are briefly reviewed; the corresponding implementation issues are discussed and shown to lead to several problems of parameter estimation. Simple situations where such constrained problems naturally arise, are presented in the context of queueing systems, in order to illustrate various points of the theory. In each case, the structure of the optimal policy is exhibited.