Stable Markov decision processes using simulation based predictive control

Stable Markov decision processes using simulation based predictive control
复制标题

DOI:
--
复制
发表时间:
2010-07
期刊:
--
影响因子:
--
通讯作者:
Zhenyao Yang;N. Kantas;Andrea Lecchini-Visintini;J. Maciejowski
Zhenyao Yang;N. Kantas;Andrea Lecchini-Visintini;J. Maciejowski
中科院分区:
其他
文献类型:
--
作者:
Zhenyao Yang;N. Kantas;Andrea Lecchini-Visintini;J. Maciejowski

文献摘要

相似文献

本文研究了在弱假设条件下模型预测控制在马尔可夫决策过程中的应用。我们基于一类特定的代价函数的最优性给出了稳定性的条件。这些结果从理论和计算的角度来看都是有用的。当考虑一般状态空间的非线性非高斯模型时,由于缺乏分析工具,因此必须使用基于仿真的方法。流行的基于仿真的方法,如随机规划和马尔可夫链蒙特卡罗,可以用来提供优化器的开环估计。考虑到这一点,我们提供了这样的方法将产生稳定的马尔可夫决策过程的条件。
In this paper we investigate the use of Model Predic- tive control for Markov Decision Processes under weak assump- tions. We provide conditions for stability based on optimality of a specific class of cost functions. These results are useful from both a theoretical and computational perspective. When nonlinear non-Gaussian models for general state spaces are considered, the absence of analytical tools makes the use of simulation based methods necessary. Popular simulation based methods like stochas- tic programming and Markov Chain Monte Carlo can be used to provide open loop estimates of the optimisers. With this in mind we provide conditions under which such an approach would yield stable Markov Decision Processes.