Iterated Deep Reinforcement Learning in Games: History-Aware Training for Improved Stability

Iterated Deep Reinforcement Learning in Games: History-Aware Training for Improved Stability
复制标题

游戏中的迭代深度强化学习:历史感知训练以提高稳定性

DOI:
10.1145/3328526.3329634
复制
发表时间:
2019
期刊:
20th ACM Conference on Economics and Computation
影响因子:
--
通讯作者:
Wellman, Michael P.
Wellman, Michael P.
中科院分区:
--
文献类型:
--
作者:
Wright, Mason;Wang, Yongzhao;Wellman, Michael P.

文献摘要

参考文献

被引文献

相似文献

深度强化学习(RL)是在复杂环境中生成策略的一种强大方法,最近在游戏方面的突破利用了深度强化学习作为迭代多代理搜索过程的一部分。我们在这些发展的基础上,提出了一种方法,通过迭代使用经验博弈论分析(EGTA)和深度RL策略,在不完全信息的复杂动态博弈中学习逐渐更好的混合策略。我们将该方法应用于定义在攻击图上的具有挑战性的网络安全游戏。用EGTA迭代深度RL到数十轮收敛,我们生成的混合策略比以前公布的启发式策略要强得多。我们通过在包括失衡但最近看到的对手的训练环境中进行微调,进一步完善了战略探索过程。实验表明,这种有历史意识的方法在训练的每个阶段都能产生较低悔恨的策略。
Deep reinforcement learning (RL) is a powerful method for generating policies in complex environments, and recent breakthroughs in game-playing have leveraged deep RL as part of an iterative multiagent search process. We build on such developments and present an approach that learns progressively better mixed strategies in complex dynamic games of imperfect information, through iterated use of empirical game-theoretic analysis (EGTA) with deep RL policies. We apply the approach to a challenging cybersecurity game defined over attack graphs. Iterating deep RL with EGTA to convergence over dozens of rounds, we generate mixed strategies far stronger than earlier published heuristic strategies for this game. We further refine the strategy-exploration process, by fine-tuning in a training environment that includes out-of-equilibrium but recently seen opponents. Experiments suggest this history-aware approach yields strategies with lower regret at each stage of training.
DOI: 10.1126/science.aar6404
发表时间: 2018-12-07
期刊: SCIENCE
影响因子: 56.9
作者:
Silver, David;Hubert, Thomas;Hassabis, Demis
通讯作者: Hassabis, Demis
多阶段攻击图安全博弈:启发式策略,以及实证博弈论分析
DOI: 10.1145/3140549.3140562
发表时间: 2017
期刊: Proceedings of the 2017 Workshop on Moving Target Defense
影响因子: --
作者:
T. Nguyen;Mason Wright;Michael P. Wellman;Satinder Singh
通讯作者: Satinder Singh
评估连续双重拍卖中非自适应交易的稳定性:一种强化学习方法
DOI: --
发表时间: 2018
期刊: AAAI Workshops
影响因子: --
作者:
Mason Wright;Michael P. Wellman
通讯作者: Michael P. Wellman
DOI: --
发表时间: 2015
期刊: Decision and Game Theory for Security
影响因子: --
作者:
K. Durkota;V. Lisý;B. Bosanský;Christopher Kiekintveld
通讯作者: Christopher Kiekintveld
使用强化学习在资源受限的环境中检测隐形僵尸网络
DOI: --
发表时间: 2017
期刊: MTD@CCS
影响因子: --
作者:
S. Venkatesan;Massimiliano Albanese;Ankit Shah;R. Ganesan;S. Jajodia
通讯作者: S. Jajodia