Nonstationary Continuous Time Markov Decision Processes with Discounted Criterion

Nonstationary Continuous Time Markov Decision Processes with Discounted Criterion
复制标题

DOI:
10.1006/jmaa.1993.1382
复制
发表时间:
1993-11
影响因子:
1.3
通讯作者:
Q. Hu
Q. Hu
中科院分区:
数学3区
文献类型:
--
作者:
Q. Hu

文献摘要

被引文献

相似文献

摘要 本文首先研究了带有折扣准则的非平稳连续时间马尔可夫决策过程。状态空间 S 和动作集 A ( i ) 是可数的,转移率 q ij ( t , a ) 和奖励率函数 r i ( t , a ) 是非齐次的。使用算子方法,我们处理该模型的最优性方程和 ϵ 最优策略的存在性。
Abstract This paper first investigates the nonstationary continuous time Markov decision processes with discounted criterion. The state space S and the action sets A ( i ) are countable, the transition rates q ij ( t , a ) and the reward rate functions r i ( t , a ) are nonhomogeneous. Using the operator method, we deal with the optimality equation and the existence of ϵ optimal policies of this model.