Deep Counterfactual Regret Minimization

Deep Counterfactual Regret Minimization
复制标题

DOI:
--
复制
发表时间:
2018-11
期刊:
--
影响因子:
--
通讯作者:
Noam Brown;Adam Lerer;Sam Gross;T. Sandholm
Noam Brown;Adam Lerer;Sam Gross;T. Sandholm
中科院分区:
其他
文献类型:
--
作者:
Noam Brown;Adam Lerer;Sam Gross;T. Sandholm

文献摘要

被引文献

相似文献

反事实遗憾最小化(CFR)是解决大型非对称信息博弈的主要框架。它通过迭代遍历博弈树收敛到一个均衡。为了处理非常大的游戏,通常在运行CFR之前应用抽象。抽象的游戏是解决与表CFR,其解决方案被映射回完整的游戏。这个过程可能是有问题的,因为抽象的方面通常是手动的和特定于领域的,抽象算法可能会错过游戏的重要战略细微差别,并且存在鸡和蛋的问题,因为确定一个好的抽象需要了解游戏的均衡。本文介绍了深度反事实遗憾最小化,这是CFR的一种形式,通过使用深度神经网络来近似CFR在整个游戏中的行为,从而避免了抽象的需要。我们表明Deep CFR是有原则的,并且在大型扑克游戏中取得了强劲的表现。这是CFR的第一个非表格变体在大型游戏中获得成功。
Counterfactual Regret Minimization (CFR) is the leading framework for solving large imperfect-information games. It converges to an equilibrium by iteratively traversing the game tree. In order to deal with extremely large games, abstraction is typically applied before running CFR. The abstracted game is solved with tabular CFR, and its solution is mapped back to the full game. This process can be problematic because aspects of abstraction are often manual and domain specific, abstraction algorithms may miss important strategic nuances of the game, and there is a chicken-and-egg problem because determining a good abstraction requires knowledge of the equilibrium of the game. This paper introduces Deep Counterfactual Regret Minimization, a form of CFR that obviates the need for abstraction by instead using deep neural networks to approximate the behavior of CFR in the full game. We show that Deep CFR is principled and achieves strong performance in large poker games. This is the first non-tabular variant of CFR to be successful in large games.