DeepFP for Finding Nash Equilibrium in Continuous Action Spaces

DeepFP for Finding Nash Equilibrium in Continuous Action Spaces
复制标题

DOI:
10.1007/978-3-030-32430-8_15
复制
发表时间:
2019-10
期刊:
--
影响因子:
--
通讯作者:
Nitin Kamra;Umang Gupta;Kai Wang;Fei Fang;Yan Liu;Milind Tambe
Nitin Kamra;Umang Gupta;Kai Wang;Fei Fang;Yan Liu;Milind Tambe
中科院分区:
其他
文献类型:
--
作者:
Nitin Kamra;Umang Gupta;Kai Wang;Fei Fang;Yan Liu;Milind Tambe

文献摘要

被引文献

相似文献

在连续行动空间中寻找纳什均衡是一个具有挑战性的问题,并且在保护地理区域免受潜在攻击等领域具有应用。我们提出了DeepFP,这是连续动作空间中虚拟游戏的近似扩展。DeepFP通过生成神经网络代表玩家的近似最佳反应,生成神经网络是高度表达的隐式密度近似器。此外,它还使用了一个游戏模型网络,该网络近似于玩家在给定行为下的预期收益,并在基于模型的学习机制中对网络进行端到端的训练。此外,如果可用,DeepFP允许使用特定领域的预言器,因此可以利用数学编程等技术来计算结构化游戏的最佳响应。我们在几个经典博弈中证明了对纳什均衡的稳定收敛,并将DeepFP应用于大型森林安全领域,并提出了一种新的防御者最佳响应预测。我们表明,DeepFP学习的策略对对抗性开发具有鲁棒性,并且随着玩家资源数量的增加而扩展得很好。
Finding Nash equilibrium in continuous action spaces is a challenging problem and has applications in domains such as protecting geographic areas from potential attackers. We present DeepFP, an approximate extension of fictitious play in continuous action spaces. DeepFP represents players’ approximate best responses via generative neural networks which are highly expressive implicit density approximators. It additionally uses a game-model network which approximates the players’ expected payoffs given their actions, and trains the networks end-to-end in a model-based learning regime. Further, DeepFP allows using domain-specific oracles if available and can hence exploit techniques such as mathematical programming to compute best responses for structured games. We demonstrate stable convergence to Nash equilibrium on several classic games and also apply DeepFP to a large forest security domain with a novel defender best response oracle. We show that DeepFP learns strategies robust to adversarial exploitation and scales well with growing number of players’ resources.