Probabilistically Guaranteed Satisfaction of Temporal Logic Constraints During Reinforcement Learning

Probabilistically Guaranteed Satisfaction of Temporal Logic Constraints During Reinforcement Learning
复制标题

强化学习期间时序逻辑约束的概率保证满足

DOI:
--
复制
发表时间:
2021
期刊:
IEEE/RJS International Conference on Intelligent RObots and Systems
影响因子:
--
通讯作者:
Ahmet Semi Asarkaya
Ahmet Semi Asarkaya
中科院分区:
--
文献类型:
--
作者:
Derya Aksaray;Yasin Yazıcıoğlu;Ahmet Semi Asarkaya

文献摘要

被引文献

相似文献

我们提出了一种新的约束强化学习方法,用于在马尔可夫决策过程中寻找最优策略,同时在整个学习过程中以期望的概率满足时间逻辑约束。提出了一种保证每一集约束概率满足的自动机理论方法,不同于在足够大的集数之后对违规行为进行惩罚以达到约束满足。该方法基于计算约束满足概率的下界,并根据需要调整勘探行为。我们给出了该方法所达到的概率约束满足的理论结果。我们还在无人机场景中数值演示了所提出的想法,其中约束是执行定期到达的拾取和交付任务,目标是飞越高奖励区域,同时执行空中监控。
We propose a novel constrained reinforcement learning method for finding optimal policies in Markov Decision Processes while satisfying temporal logic constraints with a desired probability throughout the learning process. An automata-theoretic approach is proposed to ensure the probabilistic satisfaction of the constraint in each episode, which is different from penalizing violations to achieve constraint satisfaction after a sufficiently large number of episodes. The proposed approach is based on computing a lower bound on the probability of constraint satisfaction and adjusting the exploration behavior as needed. We present theoretical results on the probabilistic constraint satisfaction achieved by the proposed approach. We also numerically demonstrate the proposed idea in a drone scenario, where the constraint is to perform periodically arriving pick-up and delivery tasks and the objective is to fly over high-reward zones to simultaneously perform aerial monitoring.