Rate of Convergence of Empirical Measures and Costs in Controlled Markov Chains and Transient Optimality

Rate of Convergence of Empirical Measures and Costs in Controlled Markov Chains and Transient Optimality
复制标题

受控马尔可夫链中经验测度和成本的收敛率及瞬态最优性

DOI:
10.1287/moor.19.4.955
复制
发表时间:
1994
期刊:
Math. Oper. Res.
影响因子:
--
通讯作者:
O. Zeitouni
O. Zeitouni
中科院分区:
--
文献类型:
--
作者:
E. Altman;O. Zeitouni

文献摘要

被引文献

相似文献

本文的目的有二。首先,在一定的递推条件下,得到了受控马氏链中经验测度收敛速度的界。这些包括通过大偏差和中心极限定理参数获得的界限。然后将这些结果应用于最优控制问题。的收敛速度的经验措施,是统一的不同的政策套上的界限,从而在收敛速度的成本上的界限。最后,新的最优控制问题,不仅涉及平均成本的标准,但也措施的瞬态行为的成本,即收敛速度,被引入并应用到一个问题,在电信。这些问题的解决方案依赖于前面几节介绍的边界。
The purpose of this paper is two fold. First, bounds on the rate of convergence of empirical measures in controlled Markov chains are obtained under some recurrence conditions. These include bounds obtained through large deviations and central limit theorem arguments. These results are then applied to optimal control problems. Bounds on the rate of convergence of the empirical measures that are uniform over different sets of policies are derived, resulting in bounds on the rate of convergence of the costs. Finally, new optimal control problems that involve not only average cost criteria but also measures on the transient behavior of the cost, namely the rate of convergence, are introduced and applied to a problem in telecommunications. The solution to these problems rely on the bounds introduced in previous sections.