Multi-Virtual-Agent Reinforcement Learning for a Stochastic Predator-Prey Grid Environment

Multi-Virtual-Agent Reinforcement Learning for a Stochastic Predator-Prey Grid Environment
复制标题

DOI:
10.1109/ijcnn55064.2022.9891898
复制
发表时间:
2022-07
期刊:
2022 International Joint Conference on Neural Networks (IJCNN)
影响因子:
--
通讯作者:
Yanbin Lin;Z. Ni;Xiangnan Zhong
Yanbin Lin;Z. Ni;Xiangnan Zhong
中科院分区:
其他
文献类型:
--
作者:
Yanbin Lin;Z. Ni;Xiangnan Zhong

文献摘要

相似文献

加强学习的概括问题至关重要,尤其对于动态环境至关重要。常规的加固学习方法通​​过一些理想的假设解决了问题,很难直接应用于动态环境。在本文中,我们为Predator-Prey网格游戏提出了一种新的多虚拟代理增强学习(MVARL)方法。即使捕食者移动,设计的方法也可以找到最佳解决方案。具体而言,我们设计虚拟代理以并行与模拟变化环境进行交互,而不是使用实际代理。此外,全球代理商从这些虚拟代理中学习信息,并同时与实际环境进行交互。这种方法不仅可以有效地改善动态环境中加强学习的概括性能,而且还可以降低整体计算成本。本文考虑了两项模拟研究,以验证该方法的有效性。我们还将结果与常规的增强学习方法进行了比较。结果表明,我们提出的方法可以改善增强学习方法的鲁棒性,并在一定程度上有助于概括。
Generalization problem of reinforcement learning is crucial especially for dynamic environments. Conventional reinforcement learning methods solve the problems with some ideal assumptions and are difficult to be applied in dynamic environments directly. In this paper, we propose a new multi-virtual- agent reinforcement learning (MVARL) approach for a predator-prey grid game. The designed method can find the optimal solution even when the predator moves. Specifically, we design virtual agents to interact with simulated changing environments in parallel instead of using actual agents. Moreover, a global agent learns information from these virtual agents and interacts with the actual environment at the same time. This method can not only effectively improve the generalization performance of reinforcement learning in dynamic environments, but also reduce the overall computational cost. Two simulation studies are considered in this paper to validate the effectiveness of the designed method. We also compare the results with the conventional reinforcement learning methods. The results indicate that our proposed method can improve the robustness of reinforcement learning method and contribute to the generalization to certain extent.