Model Learning and Knowledge Sharing for Cooperative Multiagent Systems in Stochastic Environment

Model Learning and Knowledge Sharing for Cooperative Multiagent Systems in Stochastic Environment
复制标题

DOI:
10.1109/tcyb.2019.2958912
复制
发表时间:
2021-12-01
影响因子:
11.8
通讯作者:
Li, Jr-Shin
Li, Jr-Shin
中科院分区:
计算机科学1区
文献类型:
--
作者:
Jiang, Wei-Cheng;Narayanan, Vignesh;Li, Jr-Shin

文献摘要

被引文献

相似文献

在不确定环境中,强化学习智能体面临的一项艰巨任务是快速学习一种策略或一系列动作,从而达到期望的目标。在这篇文章中,我们提出了一种增量模型学习方案来重建随机环境的模型。在所提出的学习方案中,我们引入了一个聚类算法来吸收模型信息并估计每个状态转移的概率。此外,利用重构的模型,我们提出了一种经验回放策略,通过结合探索和开发之间的平衡来创建虚拟交互体验,从而极大地加速了学习并使规划成为可能。此外,我们将所提出的学习方案扩展到多智能体框架,以减少探索所需的工作量,并减少大型环境中的学习时间。在这个多智能体框架中,我们引入了一种知识共享算法来根据需要在不同的智能体之间共享重建的模型信息,并开发了一种计算效率高的知识融合机制来融合利用智能体自身经验获得的知识和从其队友那里获得的知识。最后,给出了仿真结果并进行了对比分析,验证了所提方法在复杂学习任务中的有效性。
An imposing task for a reinforcement learning agent in an uncertain environment is to expeditiously learn a policy or a sequence of actions, with which it can achieve the desired goal. In this article, we present an incremental model learning scheme to reconstruct the model of a stochastic environment. In the proposed learning scheme, we introduce a clustering algorithm to assimilate the model information and estimate the probability for each state transition. In addition, utilizing the reconstructed model, we present an experience replay strategy to create virtual interactive experiences by incorporating a balance between exploration and exploitation, which greatly accelerates learning and enables planning. Furthermore, we extend the proposed learning scheme for a multiagent framework to decrease the effort required for exploration and to reduce the learning time in a large environment. In this multiagent framework, we introduce a knowledge-sharing algorithm to share the reconstructed model information among the different agents, as needed, and develop a computationally efficient knowledge fusing mechanism to fuse the knowledge acquired using the agents' own experience with the knowledge received from its teammates. Finally, the simulation results with comparative analysis are provided to demonstrate the efficacy of the proposed methods in the complex learning tasks.