Reinforcement Learning with Temporal Logic Constraints for Partially-Observable Markov Decision Processes
Reinforcement Learning with Temporal Logic Constraints for Partially-Observable Markov Decision Processes
复制标题
具有时态逻辑约束的强化学习,用于部分可观察的马尔可夫决策过程
DOI:
--
复制
发表时间:
2021
期刊:
影响因子:
--
通讯作者:
Miroslav Pajic
中科院分区:
文献类型:
--
作者:
Yu Wang;A. Bozkurt;Miroslav Pajic
This paper proposes a reinforcement learning method for controller synthesis of autonomous systems in unknown and partially-observable environments with subjective time-dependent safety constraints. Mathematically, we model the system dynamics by a partially-observable Markov decision process (POMDP) with unknown transition/observation probabilities. The time-dependent safety constraint is captured by iLTL, a variation of linear temporal logic for state distributions. Our Reinforcement learning method first constructs the belief MDP of the POMDP, capturing the time evolution of estimated state distributions. Then, by building the product belief MDP of the belief MDP and the limiting deterministic B\uchi automaton (LDBA) of the temporal logic constraint, we transform the time-dependent safety constraint on the POMDP into a state-dependent constraint on the product belief MDP. Finally, we learn the optimal policy by value iteration under the state-dependent constraint.
DOI:
10.1109/icra48506.2021.9561989
发表时间:
2021
期刊:
2021 IEEE International Conference on Robotics and Automation (ICRA
影响因子:
--
作者:
Bozkurt, Alper Kamil;Wang, Yu;Zavlanos, Michael M.;Pajic, Miroslav
通讯作者:
Pajic, Miroslav
DOI:
10.1109/icra48506.2021.9560940
发表时间:
2021
期刊:
2021 IEEE International Conference on Robotics and Automation (ICRA
影响因子:
--
作者:
Bozkurt, Alper Kamil;Wang, Yu;Pajic, Miroslav
通讯作者:
Pajic, Miroslav