Constrained Discounted Markov Decision Chains

Constrained Discounted Markov Decision Chains
复制标题

DOI:
10.1017/s0269964800002230
复制
发表时间:
1991-10
影响因子:
1.1
通讯作者:
L. Sennott
L. Sennott
中科院分区:
工程技术3区
文献类型:
--
作者:
L. Sennott

文献摘要

被引文献

相似文献

具有可数状态空间的马尔可夫决策链产生两种成本:运营成本和持有成本。其目标是在对预期贴现持有成本的限制下,将预期贴现运营成本降至最低。证明了最优随机简单策略的存在性。这是一种在两个固定策略之间随机化的策略,这两个策略最多在一个州不同。讨论了离散时间排队系统控制的几个例子。
A Markov decision chain with countable state space incurs two types of costs: an operating cost and a holding cost. The objective is to minimize the expected discounted operating cost, subject to a constraint on the expected discounted holding cost. The existence of an optimal randomized simple policy is proved. This is a policy that randomizes between two stationary policies, that differ in at most one state. Several examples from the control of discrete time queueing systems are discussed.