A policy iteration algorithm for Markov decision processes skip-free in one direction

A policy iteration algorithm for Markov decision processes skip-free in one direction
复制标题

DOI:
10.1145/1345263.1345360
复制
发表时间:
2007-10
期刊:
--
影响因子:
--
通讯作者:
J. Lambert;B. V. Houdt;C. Blondia
J. Lambert;B. V. Houdt;C. Blondia
中科院分区:
其他
文献类型:
--
作者:
J. Lambert;B. V. Houdt;C. Blondia

文献摘要

被引文献

相似文献

在本文中,我们提出了一种新的马尔可夫决策过程(MDP)单方向无跳过策略迭代算法。该算法基于矩阵分析方法,与 White 算法(Stochastic Models,21:785-797,2005)的精神相同,后者仅限于在两个方向上无跳跃的矩阵。当尝试提高光纤延迟线 (FDL) 缓冲器的丢失率时,可以使用马尔可夫决策过程解决的优化问题出现在光缓冲器领域。基于对此类 FDL 缓冲区的分析,我们对可用于求解 MDP 的不同技术进行了比较研究。结果表明,利用转移矩阵的结构使我们能够处理更大的系统,同时减少计算时间。
In this paper we present a new algorithm for policy iteration for Markov decision processes (MDP) skip-free in one direction. This algorithm, which is based on matrix analytic methods, is in the same spirit as the algorithm of White (Stochastic Models, 21:785-797, 2005) which was limited to matrices that are skip-free in both directions. Optimization problems that can be solved using Markov decision processes arise in the domain of optical buffers, when trying to improve loss rates of fibre delay line (FDL) buffers. Based on the analysis of such an FDL buffer we present a comparative study between the different techniques available to solve an MDP. The results illustrate that the exploitation of the structure of the transition matrices places us in a position to deal with larger systems, while reducing the computation times.