Model-Free Reinforcement Learning for Optimal Control of Markov Decision Processes Under Signal Temporal Logic Specifications

Model-Free Reinforcement Learning for Optimal Control of Markov Decision Processes Under Signal Temporal Logic Specifications
复制标题

信号时序逻辑规范下马尔可夫决策过程最优控制的无模型强化学习

DOI:
10.1109/cdc45484.2021.9683444
复制
发表时间:
2021
期刊:
2021 60th IEEE Conference on Decision and Control (CDC
影响因子:
--
通讯作者:
Nuzzo, Pierluigi
Nuzzo, Pierluigi
中科院分区:
--
文献类型:
--
作者:
Kalagarla, Krishna C.;Jain, Rahul;Nuzzo, Pierluigi

文献摘要

参考文献

被引文献

相似文献

我们提出了一种无模型强化学习(RL)算法,以找到一个最佳的政策,为有限时域马尔可夫决策过程(MDP),同时保证所需的下界的概率满足信号时序逻辑(STL)规范。我们提出了一种方法来有效地增加MDP状态空间,以捕获所需的状态历史,并表示STL目标作为一个可达性目标。规划问题,然后可以制定为一个有限时域约束马尔可夫决策过程(CMDP)。对于具有未知转移概率的一般有限时域CMDP问题,我们开发了一种强化学习方案,该方案可以利用任何无模型RL算法来提供非平稳随机策略的一般空间中的近似最优策略。我们说明了我们的方法的背景下,机器人运动规划的复杂任务的不确定性和性能目标。
We present a model-free reinforcement learning (RL) algorithm to find an optimal policy for a finite-horizon Markov decision process (MDP) while guaranteeing a desired lower bound on the probability of satisfying a signal temporal logic (STL) specification. We propose a method to effectively augment the MDP state space to capture the required state history and express the STL objective as a reachability objective. The planning problem can then be formulated as a finite-horizon constrained Markov decision process (CMDP). For a general finite-horizon CMDP problem with unknown transition probability, we develop a reinforcement learning scheme that can leverage any model-free RL algorithm to provide an approximately optimal policy out of the general space of non-stationary randomized policies. We illustrate our approach in the context of robotic motion planning for complex missions under uncertainty and performance objectives.
DOI: 10.1609/aaai.v35i9.16979
发表时间: 2020-09
期刊: ArXiv
影响因子: --
作者:
K. C. Kalagarla;Rahul Jain;P. Nuzzo
通讯作者: K. C. Kalagarla;Rahul Jain;P. Nuzzo
信号时序逻辑目标的易于处理的强化学习
DOI: --
发表时间: 2020
期刊: Conference on Learning for Dynamics & Control
影响因子: --
作者:
Harish K. Venkataraman;Derya Aksaray;P. Seiler
通讯作者: P. Seiler