Small Nash Equilibrium Certificates in Very Large Games

Small Nash Equilibrium Certificates in Very Large Games
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
ArXiv
影响因子:
--
通讯作者:
B. Zhang;T. Sandholm
B. Zhang;T. Sandholm
中科院分区:
其他
文献类型:
--
作者:
B. Zhang;T. Sandholm

文献摘要

被引文献

相似文献

在许多游戏设置中,游戏没有明确给出,而只能通过玩它来访问。虽然在这样的设置中有令人印象深刻的演示,但现有技术没有提供安全保证,即保证计算策略的博弈论可利用性。在本文中,我们介绍了一种方法,表明它是可能的,在这样的设置提供可利用性保证,而无需探索整个游戏。我们引入一个概念的一个证书的一个扩展形式的近似纳什均衡。为了验证证书,我们给出了一个算法,该算法在时间上与证书的大小成线性关系,而不是与整个博弈的大小成线性关系。在零和博弈中,我们进一步证明了最优证书-考虑到目前为止的探索-可以用任何标准的博弈求解算法(例如,使用线性程序或反事实后悔最小化)。然而,与正常形式或完美信息的情况下,我们表明,某些家庭的广泛形式的游戏没有小的近似证书,即使在非常好的假设的游戏结构。尽管有这种困难,我们发现实验非常小的证书,甚至确切的,往往存在于大型,甚至在无限的游戏。总的来说,我们的方法使人们能够尝试自己最喜欢的探索策略,同时提供可利用性保证,从而将探索策略与均衡发现过程脱钩。
In many game settings, the game is not explicitly given but is only accessible by playing it. While there have been impressive demonstrations in such settings, prior techniques have not offered safety guarantees, that is, guarantees on the game-theoretic exploitability of the computed strategies. In this paper we introduce an approach that shows that it is possible to provide exploitability guarantees in such settings without ever exploring the entire game. We introduce a notion of a certificatae of an extensive-form approximate Nash equilibrium. For verifying a certificate, we give an algorithm that runs in time linear in the size of the certificate rather than the size of the whole game. In zero-sum games, we further show that an optimal certificate---given the exploration so far---can be computed with any standard game-solving algorithm (e.g., using a linear program or counterfactual regret minimization). However, unlike in the cases of normal form or perfect information, we show that certain families of extensive-form games do not have small approximate certificates, even after making extremely nice assumptions on the structure of the game. Despite this difficulty, we find experimentally that very small certificates, even exact ones, often exist in large and even in infinite games. Overall, our approach enables one to try one's favorite exploration strategies while offering exploitability guarantees, thereby decoupling the exploration strategy from the equilibrium-finding process.