Model-Free Reinforcement Learning for Optimal Control of Markov Decision Processes Under Signal Temporal Logic Specifications
Model-Free Reinforcement Learning for Optimal Control of Markov Decision Processes Under Signal Temporal Logic Specifications
复制标题
信号时序逻辑规范下马尔可夫决策过程最优控制的无模型强化学习
DOI:
10.1109/cdc45484.2021.9683444
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Nuzzo, Pierluigi
中科院分区:
文献类型:
--
作者:
Kalagarla, Krishna C.;Jain, Rahul;Nuzzo, Pierluigi
We present a model-free reinforcement learning (RL) algorithm to find an optimal policy for a finite-horizon Markov decision process (MDP) while guaranteeing a desired lower bound on the probability of satisfying a signal temporal logic (STL) specification. We propose a method to effectively augment the MDP state space to capture the required state history and express the STL objective as a reachability objective. The planning problem can then be formulated as a finite-horizon constrained Markov decision process (CMDP). For a general finite-horizon CMDP problem with unknown transition probability, we develop a reinforcement learning scheme that can leverage any model-free RL algorithm to provide an approximately optimal policy out of the general space of non-stationary randomized policies. We illustrate our approach in the context of robotic motion planning for complex missions under uncertainty and performance objectives.
DOI:
10.1609/aaai.v35i9.16979
发表时间:
2020-09
期刊:
ArXiv
影响因子:
--
作者:
K. C. Kalagarla;Rahul Jain;P. Nuzzo
通讯作者:
K. C. Kalagarla;Rahul Jain;P. Nuzzo
DOI:
--
发表时间:
2020
期刊:
Conference on Learning for Dynamics & Control
影响因子:
--
作者:
Harish K. Venkataraman;Derya Aksaray;P. Seiler
通讯作者:
P. Seiler