Human-level performance in 3D multiplayer games with population-based reinforcement learning

Human-level performance in 3D multiplayer games with population-based reinforcement learning
复制标题

DOI:
10.1126/science.aau6249
复制
发表时间:
2019-05-31
期刊:
影响因子:
56.9
通讯作者:
Graepel, Thore
Graepel, Thore
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Jaderberg, Max;Czarnecki, Wojciech M.;Graepel, Thore

文献摘要

被引文献

相似文献

强化学习(RL)在日益复杂的单智能体环境和双人回合制游戏中取得了巨大的成功。然而,真实的世界包含多个智能体,每个智能体都独立地学习和行动,与其他智能体合作和竞争。我们使用了一个锦标赛风格的评估,以证明一个代理可以实现人类水平的性能,在三维多人第一人称视频游戏,雷神之锤III竞技场在夺旗模式,仅使用像素和游戏点得分作为输入。我们使用了一个双层优化过程,在这个过程中,一群独立的RL代理从随机生成的环境中的数千个并行匹配中同时训练。每个智能体学习自己的内部奖励信号和丰富的世界表征。这些结果表明,多智能体强化学习的人工智能研究的巨大潜力。
Reinforcement learning (RL) has shown great success in increasingly complex single-agent environments and two-player turn-based games. However, the real world contains multiple agents, each learning and acting independently to cooperate and compete with other agents. We used a tournament-style evaluation to demonstrate that an agent can achieve human-level performance in a three-dimensional multiplayer first-person video game, Quake III Arena in Capture the Flag mode, using only pixels and game points scored as input. We used a two-tier optimization process in which a population of independent RL agents are trained concurrently from thousands of parallel matches on randomly generated environments. Each agent learns its own internal reward signal and rich representation of the world. These results indicate the great potential of multiagent reinforcement learning for artificial intelligence research.