XDO: A Double Oracle Algorithm for Extensive-Form Games

XDO: A Double Oracle Algorithm for Extensive-Form Games
复制标题

DOI:
--
复制
发表时间:
2021-03
期刊:
ArXiv
影响因子:
--
通讯作者:
S. McAleer;John Lanier;P. Baldi;Roy Fox
S. McAleer;John Lanier;P. Baldi;Roy Fox
中科院分区:
其他
文献类型:
--
作者:
S. McAleer;John Lanier;P. Baldi;Roy Fox

文献摘要

被引文献

相似文献

策略空间响应预言机(PSRO)是一种用于两人零和博弈的强化学习(RL)算法,已被经验证明可以在大型博弈中找到近似的纳什均衡。虽然PSRO保证收敛到近似的纳什均衡,并且可以处理连续的动作,但随着信息状态(infostates)数量的增长,它可能需要指数数量的迭代。我们提出了扩展形式的双预言机(XDO),一个扩展形式的双预言机算法的两个球员的零和游戏,保证收敛到一个近似的纳什均衡线性的信息状态的数量。与PSRO不同,PSRO在游戏的根源混合最佳对策,XDO在每个信息状态混合最佳对策。我们还介绍了神经XDO(NXDO),其中通过深度RL学习最佳响应。在Leduc扑克的表格实验中,我们发现XDO在比PSRO小一个数量级的迭代次数中达到近似纳什均衡。在一个改进的Leduc扑克游戏和Oshi-Zumo上的实验表明,在相同的计算量下,表格XDO比CFR具有更低的可利用性。我们还发现,NXDO优于PSRO和NFSP的顺序多维连续动作游戏。NXDO是第一个可以在高维连续动作序列游戏中找到近似纳什均衡的深度RL方法。实验代码可在https://github.com/indylab/nxdo获得。
Policy Space Response Oracles (PSRO) is a reinforcement learning (RL) algorithm for two-player zero-sum games that has been empirically shown to find approximate Nash equilibria in large games. Although PSRO is guaranteed to converge to an approximate Nash equilibrium and can handle continuous actions, it may take an exponential number of iterations as the number of information states (infostates) grows. We propose Extensive-Form Double Oracle (XDO), an extensive-form double oracle algorithm for two-player zero-sum games that is guaranteed to converge to an approximate Nash equilibrium linearly in the number of infostates. Unlike PSRO, which mixes best responses at the root of the game, XDO mixes best responses at every infostate. We also introduce Neural XDO (NXDO), where the best response is learned through deep RL. In tabular experiments on Leduc poker, we find that XDO achieves an approximate Nash equilibrium in a number of iterations an order of magnitude smaller than PSRO. Experiments on a modified Leduc poker game and Oshi-Zumo show that tabular XDO achieves a lower exploitability than CFR with the same amount of computation. We also find that NXDO outperforms PSRO and NFSP on a sequential multidimensional continuous-action game. NXDO is the first deep RL method that can find an approximate Nash equilibrium in high-dimensional continuous-action sequential games. Experiment code is available at https://github.com/indylab/nxdo.