Stopped Markov decision processes with multiple constraints
Stopped Markov decision processes with multiple constraints
复制标题
停止具有多重约束的马尔可夫决策过程
DOI:
10.1007/s001860100160
复制
发表时间:
2001
影响因子:
1.2
通讯作者:
M. Horiguchi
中科院分区:
文献类型:
--
作者:
M. Horiguchi
In this paper, a optimization problem for stopped Markov decision processes with vector-valued terminal reward and multiple running cost constraints is considered. Applying the idea of occupation measures and using the scalarization technique for vector maximization problems we obtain the equivalent Mathematical Programming problem and show the existence of a Pareto optimal pair of stationary policy and stopping time requiring randomization in at most k states, where k is the number of constraints. Moreover Lagrange multiplier approaches are considered. The saddle-point statements are given, whose results are applied to obtain a related parametric Mathematical Programming, by which the problem is solved. Numerical examples are given.