Large Scale Learning of Agent Rationality in Two-Player Zero-Sum Games

Large Scale Learning of Agent Rationality in Two-Player Zero-Sum Games
复制标题

两人零和博弈中代理理性的大规模学习

DOI:
10.1609/aaai.v33i01.33016104
复制
发表时间:
2019
影响因子:
3.9
通讯作者:
J. Z. Kolter
J. Z. Kolter
中科院分区:
工程技术2区
文献类型:
--
作者:
Chun Kai Ling;Fei Fang;J. Z. Kolter

文献摘要

被引文献

相似文献

随着最近在解决大型零和广泛形式博弈方面的进展,人们对推断仅给定代理行为的潜在博弈参数的逆问题越来越感兴趣。尽管最近的一项工作提供了一个强大的可微分端到端学习框架,该框架在深度学习框架中嵌入了一个游戏求解器,允许通过反向传播学习未知的游戏参数,但该框架在应用于有限理性人类代理和大规模问题时面临显著的局限性,导致实用性差。在本文中,我们解决了这些限制,并提出了一个适用于更实际设置的框架。首先,为了学习人类智能体在复杂的二人零和博弈中的合理性,我们借鉴了决策理论中众所周知的思想,获得了一个简洁且可解释的智能体行为模型,并推导了端到端学习的求解器和梯度。其次,为了扩展到大型的现实世界场景,我们提出了一种有效的一阶原始对偶方法,该方法利用了广泛形式博弈的结构,在博弈求解和梯度计算方面产生了显着更快的计算速度。当我们在随机生成的游戏中进行测试时,我们报告的速度比之前的方法提高了几个数量级。我们还证明了我们的模型在真实世界的单人玩家设置和合成数据上的有效性。
With the recent advances in solving large, zero-sum extensive form games, there is a growing interest in the inverse problem of inferring underlying game parameters given only access to agent actions. Although a recent work provides a powerful differentiable end-to-end learning frameworks which embed a game solver within a deep-learning framework, allowing unknown game parameters to be learned via backpropagation, this framework faces significant limitations when applied to boundedly rational human agents and large scale problems, leading to poor practicality. In this paper, we address these limitations and propose a framework that is applicable for more practical settings. First, seeking to learn the rationality of human agents in complex two-player zero-sum games, we draw upon well-known ideas in decision theory to obtain a concise and interpretable agent behavior model, and derive solvers and gradients for end-to-end learning. Second, to scale up to large, real-world scenarios, we propose an efficient first-order primal-dual method which exploits the structure of extensive-form games, yielding significantly faster computation for both game solving and gradient computation. When tested on randomly generated games, we report speedups of orders of magnitude over previous approaches. We also demonstrate the effectiveness of our model on both real-world one-player settings and synthetic data.