Distributed Fictitious Play in Potential Games with Time Varying Communication Networks

Distributed Fictitious Play in Potential Games with Time Varying Communication Networks
复制标题

具有时变通信网络的潜在博弈中的分布式虚拟游戏

DOI:
10.1109/ieeeconf44664.2019.9048896
复制
发表时间:
2019
期刊:
2019 53rd Asilomar Conference on Signals, Systems, and Computers
影响因子:
--
通讯作者:
Ceyhun Eksin
Ceyhun Eksin
中科院分区:
--
文献类型:
--
作者:
Sina Arefizadeh;Ceyhun Eksin

文献摘要

被引文献

相似文献

我们提出了一个分布式算法的多智能体系统,旨在优化一个共同的目标时,代理不同的目标相关的环境状态的估计。每个代理保持对环境的估计和其他代理的行为模型。其他代理人的行为模型假设代理人根据过去行为的经验频率确定的静态分布随机选择他们的行为。在每一步中,每个智能体都采取行动,最大化其对共同目标的期望,该共同目标是根据其对环境的估计和对他人的模型计算的。我们提出了一个加权平均规则与非双重随机权重的代理估计经验频率的所有其他代理的过去的行动交换他们的估计与他们的邻居在一个随时间变化的通信网络。在这种平均规则下,我们表明代理人的估计收敛到实际的经验频率足够快。这意味着行动收敛到纳什均衡的游戏与相同的回报给出的期望共同的目标相对于一个渐近商定的估计的环境状态。
We propose a distributed algorithm for multiagent systems that aim to optimize a common objective when agents differ in their estimates of the objective-relevant state of the environment. Each agent keeps an estimate of the environment and a model of the behavior of other agents. The model of other agents’ behavior assumes agents choose their actions randomly based on a stationary distribution determined by the empirical frequencies of past actions. At each step, each agent takes the action that maximizes its expectation of the common objective computed with respect to its estimate of the environment and its model of others. We propose a weighted averaging rule with non-doubly stochastic weights for agents to estimate the empirical frequency of past actions of all other agents by exchanging their estimates with their neighbors over a time-varying communication network. Under this averaging rule, we show agents’ estimates converge to the actual empirical frequencies fast enough. This implies convergence of actions to a Nash equilibrium of the game with identical payoffs given by the expectation of the common objective with respect to an asymptotically agreed estimate of the state of the environment.