Stopped Markov decision processes with multiple constraints

Stopped Markov decision processes with multiple constraints
复制标题

停止具有多重约束的马尔可夫决策过程

DOI:
10.1007/s001860100160
复制
发表时间:
2001
影响因子:
1.2
通讯作者:
M. Horiguchi
M. Horiguchi
中科院分区:
数学4区
文献类型:
--
作者:
M. Horiguchi

文献摘要

被引文献

相似文献

研究了具有向量值终端报酬和多运行费用约束的停止马氏决策过程的优化问题。应用占用措施的思想和向量最大化问题的标量化技术,我们得到等价的数学规划问题,并显示存在一个Pareto最优的对平稳的政策和停止时间要求随机在最多k个状态,其中k是约束的数量。此外,考虑了拉格朗日乘数法。给出了鞍点的说明,并将其结果应用于一个相关的参数数学规划,从而解决了该问题。给出了数值算例。
In this paper, a optimization problem for stopped Markov decision processes with vector-valued terminal reward and multiple running cost constraints is considered. Applying the idea of occupation measures and using the scalarization technique for vector maximization problems we obtain the equivalent Mathematical Programming problem and show the existence of a Pareto optimal pair of stationary policy and stopping time requiring randomization in at most k states, where k is the number of constraints. Moreover Lagrange multiplier approaches are considered. The saddle-point statements are given, whose results are applied to obtain a related parametric Mathematical Programming, by which the problem is solved. Numerical examples are given.