Penalty-Regulated Dynamics and Robust Learning Procedures in Games

Penalty-Regulated Dynamics and Robust Learning Procedures in Games
复制标题

游戏中的惩罚调节动力学和鲁棒学习程序

DOI:
10.1287/moor.2014.0687
复制
发表时间:
2013
期刊:
Math. Oper. Res.
影响因子:
--
通讯作者:
P. Mertikopoulos
P. Mertikopoulos
中科院分区:
--
文献类型:
--
作者:
Pierre Coucheney;B. Gaujal;P. Mertikopoulos

文献摘要

被引文献

相似文献

从战略性n个游戏游戏的启发式学习方案开始,我们得出了一个新的连续时间学习动态,该动态由像复制器一样的漂移组成,并通过惩罚术语调整了,从而使游戏策略空间复制的边界。受管制的动态等同于玩家保持其正在进行的收益的指数折扣总额,然后使用平稳的最佳响应来根据这些性能得分选择动作。由于这种继承二元性,拟议的动态满足了进化游戏理论的民间理论的变体,它们会融合到(任意精确的)NASH等效游戏中的近似值。一种基于回报的学习算法,保留这些融合属性,并且只要求玩家观察其游戏中的回报。在存在随机扰动和观察错误的情况下,强大的稳定性,并且不需要玩家之间的任何同步。
Starting from a heuristic learning scheme for strategic N -person games, we derive a new class of continuous-time learning dynamics consisting of a replicator-like drift adjusted by a penalty term that renders the boundary of the game’s strategy space repelling. These penalty-regulated dynamics are equivalent to players keeping an exponentially discounted aggregate of their ongoing payoffs and then using a smooth best response to pick an action based on these performance scores. Owing to this inherent duality, the proposed dynamics satisfy a variant of the folk theorem of evolutionary game theory and they converge to (arbitrarily precise) approximations of Nash equilibria in potential games. Motivated by applications to traffic engineering, we exploit this duality further to design a discrete-time, payoff-based learning algorithm that retains these convergence properties and only requires players to observe their in-game payoffs. Moreover, the algorithm remains robust in the presence of stochastic perturbations and observation errors, and it does not require any synchronization between players.