A composite learning method for multi-ship collision avoidance based on reinforcement learning and inverse control

A composite learning method for multi-ship collision avoidance based on reinforcement learning and inverse control
复制标题

基于强化学习和逆控制的多船避碰复合学习方法

DOI:
10.1016/j.neucom.2020.05.089
复制
发表时间:
2020-10-21
期刊:
影响因子:
6
通讯作者:
Liu, Chenguang
Liu, Chenguang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Xie, Shuo;Chu, Xiumin;Liu, Chenguang

文献摘要

被引文献

相似文献

无模型强化学习方法在未知环境下船舶避碰方面具有潜力。针对无模型强化学习效率低的问题,提出了一种基于异步优势actor-critic (A3C)算法、长短期记忆神经网络(LSTM)和q学习的复合学习方法。该方法使用q学习在基于LSTM逆模型的控制器和无模型A3C策略之间进行自适应决策。通过多船避碰仿真,验证了无模型A3C方法、基于逆模型的方法和复合学习方法的有效性。仿真结果表明,基于复合学习的船舶避碰方法优于A3C学习方法和传统的基于优化的避碰方法。(c) 2020 Elsevier B.V.版权所有
Model-free reinforcement learning methods have potentials in ship collision avoidance under unknown environments. To defect the low efficiency problem of the model-free reinforcement learning, a composite learning method is proposed based on an asynchronous advantage actor-critic (A3C) algorithm, a long short-term memory neural network (LSTM) and Q-learning. The proposed method uses Q-learning for adaptive decisions between a LSTM inverse model-based controller and the model-free A3C policy. Multi-ship collision avoidance simulations are conducted to verify the effectiveness of the model-free A3C method, the proposed inverse model-based method and the composite learning method. The simulation results indicate that the proposed composite learning based ship collision avoidance method outperforms the A3C learning method and a traditional optimization-based method. (c) 2020 Elsevier B.V. All rights reserved.