Decentralized inertial best-response with voluntary and limited communication in random communication networks

Decentralized inertial best-response with voluntary and limited communication in random communication networks
复制标题

随机通信网络中自愿和有限通信的去中心化惯性最佳响应

DOI:
10.1016/j.automatica.2022.110566
复制
发表时间:
2022
期刊:
影响因子:
6.4
通讯作者:
Eksin, Ceyhun
Eksin, Ceyhun
中科院分区:
计算机科学2区
文献类型:
--
作者:
Aydın, Sarper;Eksin, Ceyhun

文献摘要

相似文献

多个自主代理通过随机通信网络进行交互,以最大化其依赖于其他代理的行为的个体效用函数。我们考虑分散的最佳反应与惯性型算法,在该算法中,代理形成信念的其他球员的未来行动的基础上,本地信息,并采取行动,最大限度地提高他们的预期效用计算这些信念或继续采取他们以前的行动。我们证明了这些类型的算法收敛到弱非循环博弈的纳什均衡。结果取决于信念更新和信息交换协议在给定静态环境的有限时间内以正概率成功学习其他参与者的动作的条件,即,当其他代理人的行为没有改变时。我们设计了一个分散的虚拟播放算法与自愿和有限的通信(DFP-VL)协议,满足这个条件。在自愿通信协议中,每个智能体通过评估其信息的新奇和其信息对其他人的信念的潜在影响来决定与谁交换信息。有限的通信协议需要代理只发送他们最频繁的动作,他们决定与代理进行通信。目标分配游戏的数值实验表明,自愿和有限的通信协议可以一半以上的通信尝试的数量,同时保持相同的收敛速度DFP在代理不断尝试进行通信。
Multiple autonomous agents interact over a random communication network to maximize their individual utility functions which depend on the actions of other agents. We consider decentralized best-response with inertia type algorithms in which agents form beliefs about the future actions of other players based on local information, and take actions that maximize their expected utilities computed with respect to these beliefs or continue to take their previous actions. We show convergence of these types of algorithms to a Nash equilibrium in weakly acyclic games. The result depends on the condition that the belief update and information exchange protocols successfully learn the actions of other players with positive probability in finite time given a static environment, i.e., when other agents’ actions do not change. We design a decentralized fictitious play algorithm with voluntary and limited communication (DFP-VL) protocols that satisfy this condition. In the voluntary communication protocol, each agent decides whom to exchange information with by assessing the novelty of its information and the potential effect of its information on others’ beliefs. The limited communication protocol entails agents sending only their most frequent action to agents that they decide to communicate with. Numerical experiments on a target assignment game demonstrate that the voluntary and limited communication protocol can more than halve the number of communication attempts while retaining the same convergence rate as DFP in which agents constantly attempt to communicate.