Learning to Play Pursuit-Evasion with Visibility Constraints

Learning to Play Pursuit-Evasion with Visibility Constraints
复制标题

DOI:
10.1109/iros51168.2021.9635959
复制
发表时间:
2021-09
期刊:
2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Selim Engin;Qingyuan Jiang;Volkan Isler
Selim Engin;Qingyuan Jiang;Volkan Isler
中科院分区:
其他
文献类型:
--
作者:
Selim Engin;Qingyuan Jiang;Volkan Isler

文献摘要

相似文献

我们研究了多边形环境中单个追逐者和一个逃避者的追逐-逃避问题,其中玩家具有可见性约束。追捕者的任务是尽快抓住逃避者,而逃避者则试图避免被抓获。我们把这个问题形式化为一个零和博弈,在这个博弈中,参与者有私人的观察和相互冲突的目标。例如,当一个玩家,追捕者没有看到逃避者时,它需要推理逃避者的所有可能位置。与竞技场大小相比,这导致状态空间的大小呈指数增加。为了克服与大状态空间相关的挑战,我们引入了一种新的基于学习的方法,该方法压缩游戏状态并使用它来为玩家规划行动。结果表明,我们的方法优于现有的强化学习方法,并在复杂环境中与当前最先进的随机策略竞争。
We study the problem of pursuit-evasion for a single pursuer and an evader in polygonal environments where the players have visibility constraints. The pursuer is tasked with catching the evader as quickly as possible while the evader tries to avoid being captured. We formalize this problem as a zero-sum game where the players have private observations and conflicting objectives.One of the challenging aspects of this game is due to limited visibility. When a player, for example, the pursuer does not see the evader, it needs to reason about all possible locations of the evader. This causes an exponential increase in the size of the state space as compared to the arena size. To overcome the challenges associated with large state spaces, we introduce a new learning-based method that compresses the game state and uses it to plan actions for the players. The results indicate that our method outperforms the existing reinforcement learning methods, and performs competitively against the current state-of-the-art randomized strategy in complex environments.