Deep Reinforcement Learning Based Networked Control with Network Delays for Signal Temporal Logic Specifications

Deep Reinforcement Learning Based Networked Control with Network Delays for Signal Temporal Logic Specifications
复制标题

DOI:
10.1109/etfa52439.2022.9921505
复制
发表时间:
2021-08
期刊:
2022 IEEE 27th International Conference on Emerging Technologies and Factory Automation (ETFA)
影响因子:
--
通讯作者:
Junya Ikemoto;T. Ushio
Junya Ikemoto;T. Ushio
中科院分区:
其他
文献类型:
--
作者:
Junya Ikemoto;T. Ushio

文献摘要

相似文献

我们将深度强化学习(DRL)应用于具有网络时延的网络控制器的设计,以完成由信号时态逻辑(STL)公式描述的时态控制任务。STL适用于处理动态系统中具有有界时间间隔的时态规范。通常,代理不仅需要当前的系统状态,还需要系统的过去行为来确定满足给定STL公式的期望控制操作。此外,我们还需要考虑网络延迟对数据传输的影响。因此,我们提出了一种利用系统过去的状态和控制动作的扩展马尔可夫决策过程,称为τd-MDP,以便智能体能够评估考虑网络时延的STL公式的满意度。在此基础上,利用τd-mdp算法设计了一个网络化控制器。通过仿真验证了该算法的学习性能。
We apply deep reinforcement learning (DRL) to design of a networked controller with network delays to complete a temporal control task that is described by a signal temporal logic (STL) formula. STL is useful to deal with a temporal specification with a bounded time interval for a dynamical system. In general, an agent needs not only the current system state but also the past behavior of the system to determine a desired control action for satisfying the given STL formula. Additionally, we need to consider the effect of network delays for data transmissions. Thus, we propose an extended Markov decision process (MDP) using past system states and control actions, which is called a τd-MDP, so that the agent can evaluate the satisfaction of the STL formula considering the network delays. Thereafter, we apply a DRL algorithm to design a networked controller using the τd-MDP. Through simulations, we also demonstrate the learning performance of the proposed algorithm.