MDP Optimal Control under Temporal Logic Constraints-Technical Report -

MDP Optimal Control under Temporal Logic Constraints-Technical Report -
复制标题

时间逻辑约束下的MDP最优控制-技术报告-

DOI:
--
复制
发表时间:
2011
期刊:
影响因子:
--
通讯作者:
D. Rus
D. Rus
中科院分区:
--
文献类型:
--
作者:
X. Ding;Stephen L. Smith;C. Belta;D. Rus

文献摘要

被引文献

相似文献

在本文中,我们开发了一种方法来自动生成一个控制策略的动态系统建模为马尔可夫决策过程(MDP)。控制规范是作为一个线性时序逻辑(LTL)公式在一组命题上定义的MDP的状态。我们合成的控制策略,使MDP满足给定的规格几乎肯定,如果这样的政策存在。此外,我们指定了一个“优化命题”,以反复满足,我们制定了一个新的优化标准,在满足这个命题之间的预期成本最小化。我们提出了一个充分条件的政策是最佳的,并开发了一个动态规划算法,合成的政策,在某些条件下是最佳的,否则次优。这个问题是由机器人应用程序需要持久的任务,如环境监测或数据收集,要执行的动机。
In this paper, we develop a method to automatically generate a control policy for a dynamical system modeled as a Markov Decision Process (MDP). The control specification is given as a Linear Temporal Logic (LTL) formula over a set of propositions defined on the states of the MDP. We synthesize a control policy such that the MDP satisfies the given specification almost surely, if such a policy exists. In addition, we designate an “optimizing proposition” to be repeatedly satisfied, and we formulate a novel optimization criterion in terms of minimizing the expected cost in between satisfactions of this proposition. We propose a sufficient condition for a policy to be optimal, and develop a dynamic programming algorithm that synthesizes a policy that is optimal under some conditions, and sub-optimal otherwise. This problem is motivated by robotic applications requiring persistent tasks, such as environmental monitoring or data gathering, to be performed.