Last-Iterate Convergence: Zero-Sum Games and Constrained Min-Max Optimization

Last-Iterate Convergence: Zero-Sum Games and Constrained Min-Max Optimization
复制标题

DOI:
10.4230/lipics.itcs.2019.27
复制
发表时间:
2018-07
期刊:
--
影响因子:
--
通讯作者:
C. Daskalakis;Ioannis Panageas
C. Daskalakis;Ioannis Panageas
中科院分区:
其他
文献类型:
--
作者:
C. Daskalakis;Ioannis Panageas

文献摘要

被引文献

相似文献

受博弈论、优化和生成对抗网络中的应用的启发,Daskalakis等人的最近工作以及Liang和Stokes的后续工作已经建立了广泛使用的梯度下降/上升过程的变体,称为“乐观梯度下降/上升(OGDA)",在{\em无约束}凹凸极小-极大优化问题中表现出最后迭代收敛到鞍点。我们表明,同样适用于更一般的问题{\em约束}的最小-最大优化下的一个变体的无悔乘法权重更新方法称为“乐观乘法权重更新(OMWU)"。这回答了Syrgkanis等人的一个未决问题。我们的结果的证明需要从根本上不同的技术,从那些存在于无悔学习文献和上述论文。我们表明,OMWU单调改善的Kullback-Leibler发散的电流的(适当规范化)最小最大的解决方案,直到它进入附近的解决方案。在该邻域内,我们表明,OMWU成为一个收缩映射收敛到精确解。我们相信,我们的技术将是有用的,在其他学习算法的最后几个月的分析。
Motivated by applications in Game Theory, Optimization, and Generative Adversarial Networks, recent work of Daskalakis et al~\cite{DISZ17} and follow-up work of Liang and Stokes~\cite{LiangS18} have established that a variant of the widely used Gradient Descent/Ascent procedure, called "Optimistic Gradient Descent/Ascent (OGDA)", exhibits last-iterate convergence to saddle points in {\em unconstrained} convex-concave min-max optimization problems. We show that the same holds true in the more general problem of {\em constrained} min-max optimization under a variant of the no-regret Multiplicative-Weights-Update method called "Optimistic Multiplicative-Weights Update (OMWU)". This answers an open question of Syrgkanis et al~\cite{SALS15}. The proof of our result requires fundamentally different techniques from those that exist in no-regret learning literature and the aforementioned papers. We show that OMWU monotonically improves the Kullback-Leibler divergence of the current iterate to the (appropriately normalized) min-max solution until it enters a neighborhood of the solution. Inside that neighborhood we show that OMWU becomes a contracting map converging to the exact solution. We believe that our techniques will be useful in the analysis of the last iterate of other learning algorithms.