Deep Reinforcement Learning Under Signal Temporal Logic Constraints Using Lagrangian Relaxation

Deep Reinforcement Learning Under Signal Temporal Logic Constraints Using Lagrangian Relaxation
复制标题

DOI:
10.1109/access.2022.3218216
复制
发表时间:
2022-01
期刊:
影响因子:
3.9
通讯作者:
Junya Ikemoto;T. Ushio
Junya Ikemoto;T. Ushio
中科院分区:
计算机科学3区
文献类型:
--
作者:
Junya Ikemoto;T. Ushio

文献摘要

相似文献

深度强化学习(Deep Reinforcement Learning,DRL)作为一种在没有系统数学模型的情况下解决最优控制问题的方法,引起了广泛的关注。另一方面,在一般情况下,约束可以施加在最优控制问题上。在本研究中,我们考虑最优控制问题的约束,以完成时间控制任务。我们使用信号时序逻辑(STL),这是有用的时间敏感的控制任务,因为它可以指定连续的信号在有限的时间间隔内的约束。为了处理STL的约束,我们引入了一个扩展的约束马尔可夫决策过程(CMDP),这被称为一个$\tau $ -CMDP。我们制定STL约束的最优控制问题的$\tau $ -CMDP,并提出了一个两阶段的约束DRL算法使用拉格朗日松弛法。通过仿真,我们也证明了所提出的算法的学习性能。
Deep reinforcement learning (DRL) has attracted much attention as an approach to solve optimal control problems without mathematical models of systems. On the other hand, in general, constraints may be imposed on optimal control problems. In this study, we consider the optimal control problems with constraints to complete temporal control tasks. We describe the constraints using signal temporal logic (STL), which is useful for time sensitive control tasks since it can specify continuous signals within bounded time intervals. To deal with the STL constraints, we introduce an extended constrained Markov decision process (CMDP), which is called a $\tau $ -CMDP. We formulate the STL-constrained optimal control problem as the $\tau $ -CMDP and propose a two-phase constrained DRL algorithm using the Lagrangian relaxation method. Through simulations, we also demonstrate the learning performance of the proposed algorithm.