Discounted Continuous-Time Controlled Markov Chains: Convergence of Control Models

Discounted Continuous-Time Controlled Markov Chains: Convergence of Control Models
复制标题

DOI:
10.1017/s0021900200012882
复制
发表时间:
2012-12
期刊:
J. Appl. Probab.
影响因子:
--
通讯作者:
T. Prieto-Rumeau;O. Hernández-Lerma
T. Prieto-Rumeau;O. Hernández-Lerma
中科院分区:
其他
文献类型:
--
作者:
T. Prieto-Rumeau;O. Hernández-Lerma

文献摘要

被引文献

相似文献

我们对连续时间、可枚举状态控制马尔可夫链(CMC)感兴趣,在折扣奖励最优性标准下,具有紧凑的 Borel 动作集,以及可能无界的转移和奖励率。对于此类 CMC,我们提出了收敛于给定控制模型 M 的控制模型序列 {Mn} 的定义,这确保了 Mn 的折扣最优奖励和策略收敛到 M 的策略。作为应用,我们提出了原始控制模型 M 的有限状态和有限动作截断技术,该技术通过对具有灾难的受控种群系统的最优奖励和策略进行数值逼近来说明。我们研究相应的收敛速度。
We are interested in continuous-time, denumerable state controlled Markov chains (CMCs), with compact Borel action sets, and possibly unbounded transition and reward rates, under the discounted reward optimality criterion. For such CMCs, we propose a definition of a sequence of control models {Mn} converging to a given control model M, which ensures that the discount optimal reward and policies of Mn converge to those of M. As an application, we propose a finite-state and finite-action truncation technique of the original control model M, which is illustrated by approximating numerically the optimal reward and policy of a controlled population system with catastrophes. We study the corresponding convergence rates.