Safety-Aware Pursuit-Evasion Games in Unknown Environments Using Gaussian Processes and Finite-Time Convergent Reinforcement Learning

Safety-Aware Pursuit-Evasion Games in Unknown Environments Using Gaussian Processes and Finite-Time Convergent Reinforcement Learning
复制标题

DOI:
10.1109/tnnls.2022.3203977
复制
发表时间:
2022-10
影响因子:
10.4
通讯作者:
Nikolaos-Marios T Kokolakis;K. Vamvoudakis
Nikolaos-Marios T Kokolakis;K. Vamvoudakis
中科院分区:
计算机科学1区
文献类型:
--
作者:
Nikolaos-Marios T Kokolakis;K. Vamvoudakis

文献摘要

被引文献

相似文献

本文开发了一个安全的追逐-逃避游戏,使有限的时间捕获,最优的性能,以及适应未知的杂乱环境。追逃博弈被描述为一个零和微分博弈,其中追赶者寻求最小化其与目标的相对距离,而逃避者试图最大化该距离。然后提出了一种基于批评者的强化学习(RL)算法,用于在线学习并在有限时间内学习追逃策略,从而实现对逃避者的有限时间捕获。通过与障碍物相关的障碍物功能来确保安全,这些功能被整合到运行成本中。利用高斯过程,设计了一种基于学习的机制来安全地学习未知环境。仿真结果表明了该方法的有效性。
This article develops a safe pursuit-evasion game for enabling finite-time capture, optimal performance as well as adaptation to an unknown cluttered environment. The pursuit-evasion game is formulated as a zero-sum differential game wherein the pursuer seeks to minimize its relative distance to the target while the evader attempts to maximize it. A critic-only reinforcement learning (RL)-based algorithm is then proposed for learning online and in finite time the pursuit-evasion policies and thus enabling finite-time capture of the evader. Safety is ensured by means of barrier functions associated with the obstacles, which are integrated into the running cost. Using Gaussian processes (GPs), a learning-based mechanism is devised for safely learning the unknown environment. Simulation results illustrate the efficacy of the proposed approach.