Consensus Multiplicative Weights Update: Learning to Learn using Projector-based Game Signatures

Consensus Multiplicative Weights Update: Learning to Learn using Projector-based Game Signatures
复制标题

共识乘法权重更新:学习使用基于投影仪的游戏签名来学习

DOI:
--
复制
发表时间:
2021
期刊:
International Conference on Machine Learning
影响因子:
--
通讯作者:
Sumitra Ganesh
Sumitra Ganesh
中科院分区:
--
文献类型:
--
作者:
N. Vadori;Rahul Savani;Thomas Spooner;Sumitra Ganesh

文献摘要

参考文献

被引文献

相似文献

Cheung 和 Piliouras (2020) 最近表明,乘法权重更新方法的两种变体 - OMWU 和 MWU - 根据游戏是零和博弈还是合作博弈而表现出相反的收敛特性。受这项工作和最近关于学习优化单一函数的文献的启发,我们引入了一种新的框架,用于学习游戏中纳什均衡的最后迭代收敛,其中更新规则沿轨迹的系数(学习率)是通过以游戏性质为条件的强化学习策略来学习的:\textit{游戏签名}。我们使用将两人游戏新分解为与交换投影算子相对应的八个组件来构建后者,从而概括和统一了文献中研究的最新游戏概念。我们比较了学习各种更新规则的系数后的性能,并表明 RL 策略能够利用各种游戏类型的游戏签名。在此过程中,我们引入了 CMWU,这是一种将共识优化扩展到约束情况的新算法,对零和双矩阵博弈具有局部收敛保证,并表明它在具有常数系数的零和博弈和学习其系数时在一系列博弈中都具有竞争性能。
Cheung and Piliouras (2020) recently showed that two variants of the Multiplicative Weights Update method - OMWU and MWU - display opposite convergence properties depending on whether the game is zero-sum or cooperative. Inspired by this work and the recent literature on learning to optimize for single functions, we introduce a new framework for learning last-iterate convergence to Nash Equilibria in games, where the update rule's coefficients (learning rates) along a trajectory are learnt by a reinforcement learning policy that is conditioned on the nature of the game: \textit{the game signature}. We construct the latter using a new decomposition of two-player games into eight components corresponding to commutative projection operators, generalizing and unifying recent game concepts studied in the literature. We compare the performance of various update rules when their coefficients are learnt, and show that the RL policy is able to exploit the game signature across a wide range of game types. In doing so, we introduce CMWU, a new algorithm that extends consensus optimization to the constrained case, has local convergence guarantees for zero-sum bimatrix games, and show that it enjoys competitive performance on both zero-sum games with constant coefficients and across a spectrum of games when its coefficients are learnt.
DOI: 10.4230/lipics.itcs.2019.27
发表时间: 2018-07
期刊: --
影响因子: --
作者:
C. Daskalakis;Ioannis Panageas
通讯作者: C. Daskalakis;Ioannis Panageas