Comprehensive cooperative deep deterministic policy gradients for multi-agent systems in unstable environment

Comprehensive cooperative deep deterministic policy gradients for multi-agent systems in unstable environment
复制标题

DOI:
10.1117/12.2519153
复制
发表时间:
2019-05
期刊:
--
影响因子:
--
通讯作者:
Dong Xie;Xiangnan Zhong;Qing Yang;Y. Huang
Dong Xie;Xiangnan Zhong;Qing Yang;Y. Huang
中科院分区:
其他
文献类型:
--
作者:
Dong Xie;Xiangnan Zhong;Qing Yang;Y. Huang

文献摘要

相似文献

如今,智能无人车辆,如无人飞机和坦克,参与了现代战场上的许多复杂任务。它们构成了具有不同程度的作战自主性的网络化智能系统,在未来战场上将继续得到越来越广泛的应用。为了应对这样一个高度不稳定的环境,智能代理需要协作来探索信息并实现整个目标。本文提出了一种新的综合协作深度确定性策略梯度(C2DDPG)算法,通过为每个Agent设计一个特殊的奖励函数来帮助协作和探索。智能体将从相邻的队友那里接收状态信息,以实现更好的团队合作。该方法在一个实时战略游戏《星际争霸》的微观管理中得到了演示,该游戏类似于一个有两组单位的战场。
Nowadays, intelligent unmanned vehicles, such as unmanned aircraft and tanks, are involved in many complex tasks in the modern battlefield. They compose the networked intelligent systems with varying degrees of operational autonomy, which will continue to be used increasingly on the future battlefield. To deal with such a highly unstable environment, intelligent agents need to collaborate to explore the information and achieve the entire goal. In this paper, we will establish a novel comprehensive cooperative deep deterministic policy gradients (C2DDPG) algorithm by designing a special reward function for each agent to help collaboration and exploration. The agents will receive states information from their neighboring teammates to achieve better teamwork. The method is demonstrated in a real-time strategy game, StarCraft micromanagement, which is similar to a battlefield with two groups of units.