Reinforcement Learning with Temporal Logic Constraints for Partially-Observable Markov Decision Processes

Reinforcement Learning with Temporal Logic Constraints for Partially-Observable Markov Decision Processes
复制标题

具有时态逻辑约束的强化学习,用于部分可观察的马尔可夫决策过程

DOI:
--
复制
发表时间:
2021
期刊:
arXiv.org
影响因子:
--
通讯作者:
Miroslav Pajic
Miroslav Pajic
中科院分区:
--
文献类型:
--
作者:
Yu Wang;A. Bozkurt;Miroslav Pajic

文献摘要

参考文献

被引文献

相似文献

本文提出了一种强化学习方法,用于在具有主观时变安全约束的未知和部分可观测环境中的自治系统的控制器综合。在数学上,我们用具有未知转移/观测概率的部分可观测马尔可夫决策过程(POMDP)来建模系统动力学。依赖于时间的安全约束被iLTL捕获,iLTL是用于状态分布的线性时态逻辑的变体。我们的强化学习方法首先构造POMDP的信念MDP,捕捉估计状态分布的时间演化。然后,通过建立信念MDP的乘积信念MDP和时态逻辑约束的极限确定性B\uchi自动机(LDBA),将POMDP上的时间相关安全约束转化为产品信念MDP上的状态相关约束。最后,在状态依赖约束下,通过数值迭代得到最优策略。
This paper proposes a reinforcement learning method for controller synthesis of autonomous systems in unknown and partially-observable environments with subjective time-dependent safety constraints. Mathematically, we model the system dynamics by a partially-observable Markov decision process (POMDP) with unknown transition/observation probabilities. The time-dependent safety constraint is captured by iLTL, a variation of linear temporal logic for state distributions. Our Reinforcement learning method first constructs the belief MDP of the POMDP, capturing the time evolution of estimated state distributions. Then, by building the product belief MDP of the belief MDP and the limiting deterministic B\uchi automaton (LDBA) of the temporal logic constraint, we transform the time-dependent safety constraint on the POMDP into a state-dependent constraint on the product belief MDP. Finally, we learn the optimal policy by value iteration under the state-dependent constraint.
具有线性时态逻辑目标的随机博弈的无模型强化学习
DOI: 10.1109/icra48506.2021.9561989
发表时间: 2021
期刊: 2021 IEEE International Conference on Robotics and Automation (ICRA
影响因子: --
作者:
Bozkurt, Alper Kamil;Wang, Yu;Zavlanos, Michael M.;Pajic, Miroslav
通讯作者: Pajic, Miroslav
通过无模型强化学习制定针对隐秘攻击的安全规划
DOI: 10.1109/icra48506.2021.9560940
发表时间: 2021
期刊: 2021 IEEE International Conference on Robotics and Automation (ICRA
影响因子: --
作者:
Bozkurt, Alper Kamil;Wang, Yu;Pajic, Miroslav
通讯作者: Pajic, Miroslav