Semicentralized Deep Deterministic Policy Gradient in Cooperative StarCraft Games

Semicentralized Deep Deterministic Policy Gradient in Cooperative StarCraft Games
复制标题

DOI:
10.1109/tnnls.2020.3042943
复制
发表时间:
2020-12
影响因子:
10.4
通讯作者:
Dong Xie;Xiangnan Zhong
Dong Xie;Xiangnan Zhong
中科院分区:
计算机科学1区
文献类型:
--
作者:
Dong Xie;Xiangnan Zhong

文献摘要

相似文献

在本文中,我们提出了一种新颖的深度确定的政策梯度(SCDDPG),以供合作游戏,我们设计了两级参与者的批评结构,以帮助代理商在当地的演员中进行互动和合作为每种代理建立了批判性结构,并从环境中收到了部分可观察的信息。本地设计基于有限的集中信息(例如健康价值)的整体视图。 ,这种设计可以通过在学习过程中向全球水平发送有限的信息来减少沟通燃烧。基于代理的属性进一步改善了随机环境中的学习性能。他们的队友在各种星际争霸场景中击败敌人。
In this article, we propose a novel semicentralized deep deterministic policy gradient (SCDDPG) algorithm for cooperative multiagent games. Specifically, we design a two-level actor-critic structure to help the agents with interactions and cooperation in the StarCraft combat. The local actor-critic structure is established for each kind of agents with partially observable information received from the environment. Then, the global actor-critic structure is built to provide the local design an overall view of the combat based on the limited centralized information, such as the health value. These two structures work together to generate the optimal control action for each agent and to achieve better cooperation in the games. Comparing with the fully centralized methods, this design can reduce the communication burden by only sending limited information to the global level during the learning process. Furthermore, the reward functions are also designed for both local and global structures based on the agents’ attributes to further improve the learning performance in the stochastic environment. The developed method has been demonstrated on several scenarios in a real-time strategy game, i.e., StarCraft. The simulation results show that the agents can effectively cooperate with their teammates and defeat the enemies in various StarCraft scenarios.