Iterated Deep Reinforcement Learning in Games: History-Aware Training for Improved Stability
Iterated Deep Reinforcement Learning in Games: History-Aware Training for Improved Stability
复制标题
游戏中的迭代深度强化学习:历史感知训练以提高稳定性
DOI:
10.1145/3328526.3329634
复制
发表时间:
2019
期刊:
影响因子:
--
通讯作者:
Wellman, Michael P.
中科院分区:
文献类型:
--
作者:
Wright, Mason;Wang, Yongzhao;Wellman, Michael P.
Deep reinforcement learning (RL) is a powerful method for generating policies in complex environments, and recent breakthroughs in game-playing have leveraged deep RL as part of an iterative multiagent search process. We build on such developments and present an approach that learns progressively better mixed strategies in complex dynamic games of imperfect information, through iterated use of empirical game-theoretic analysis (EGTA) with deep RL policies. We apply the approach to a challenging cybersecurity game defined over attack graphs. Iterating deep RL with EGTA to convergence over dozens of rounds, we generate mixed strategies far stronger than earlier published heuristic strategies for this game. We further refine the strategy-exploration process, by fine-tuning in a training environment that includes out-of-equilibrium but recently seen opponents. Experiments suggest this history-aware approach yields strategies with lower regret at each stage of training.
登录
查看更多内容
影响因子:
56.9
作者:
Silver, David;Hubert, Thomas;Hassabis, Demis
通讯作者:
Hassabis, Demis
DOI:
10.1145/3140549.3140562
发表时间:
2017
期刊:
Proceedings of the 2017 Workshop on Moving Target Defense
影响因子:
--
作者:
T. Nguyen;Mason Wright;Michael P. Wellman;Satinder Singh
通讯作者:
Satinder Singh
DOI:
--
发表时间:
2018
期刊:
AAAI Workshops
影响因子:
--
作者:
Mason Wright;Michael P. Wellman
通讯作者:
Michael P. Wellman
DOI:
--
发表时间:
2015
期刊:
Decision and Game Theory for Security
影响因子:
--
作者:
K. Durkota;V. Lisý;B. Bosanský;Christopher Kiekintveld
通讯作者:
Christopher Kiekintveld
DOI:
--
发表时间:
2017
期刊:
MTD@CCS
影响因子:
--
作者:
S. Venkatesan;Massimiliano Albanese;Ankit Shah;R. Ganesan;S. Jajodia
通讯作者:
S. Jajodia