Implicit Learning Dynamics in Stackelberg Games: Equilibria Characterization, Convergence Analysis, and Empirical Study

Implicit Learning Dynamics in Stackelberg Games: Equilibria Characterization, Convergence Analysis, and Empirical Study
复制标题

DOI:
--
复制
发表时间:
2020-07
期刊:
--
影响因子:
--
通讯作者:
Tanner Fiez;Benjamin J. Chasnov;L. Ratliff
Tanner Fiez;Benjamin J. Chasnov;L. Ratliff
中科院分区:
其他
文献类型:
--
作者:
Tanner Fiez;Benjamin J. Chasnov;L. Ratliff

文献摘要

相似文献

当代关于连续博弈学习的研究通常忽略了机器学习问题中存在的分层决策结构,而是将其视为同时进行的博弈,并采用纳什均衡解的概念。我们偏离了这种范式,并提供了一个全面的研究在Stackelberg游戏的学习。这项工作提供了深入了解零和游戏的优化景观之间建立联系纳什和Stackelberg均衡沿着与极限点的同时梯度下降。我们得到新的基于梯度的学习动态模仿自然结构的Stackelberg游戏使用隐函数定理,并提供收敛性分析确定性和随机更新的零和一般和游戏。值得注意的是,在使用确定性更新的零和游戏中,我们证明了动态收敛到的唯一临界点是Stackelberg平衡,并提供了局部收敛速度。从经验上讲,与同步梯度下降相比,我们的学习动态减轻了旋转行为,并表现出训练生成对抗网络的贝内。
Contemporary work on learning in continuous games has commonly overlooked the hierarchical decision-making structure present in machine learning problems formulated as games, instead treating them as simultaneous play games and adopting the Nash equilibrium solution concept. We deviate from this paradigm and provide a comprehensive study of learning in Stackelberg games. This work provides insights into the optimization landscape of zero-sum games by establishing connections between Nash and Stackelberg equilibria along with the limit points of simultaneous gradient descent. We derive novel gradient-based learning dynamics emulating the natural structure of a Stackelberg game using the implicit function theorem and provide convergence analysis for deterministic and stochastic updates for zero-sum and general-sum games. Notably, in zero-sum games using deterministic updates, we show the only critical points the dynamics converge to are Stackelberg equilibria and provide a local convergence rate. Empirically, our learning dynamics mitigate rotational behavior and exhibit benefits for training generative adversarial networks compared to simultaneous gradient descent.