Learning in Games with Lossy Feedback

Learning in Games with Lossy Feedback
复制标题

在有损反馈的游戏中学习

DOI:
--
复制
发表时间:
2018
期刊:
Neural Information Processing Systems
影响因子:
--
通讯作者:
Y. Ye
Y. Ye
中科院分区:
--
文献类型:
--
作者:
Zhengyuan Zhou;P. Mertikopoulos;S. Athey;N. Bambos;P. Glynn;Y. Ye

文献摘要

被引文献

相似文献

我们考虑了一个博弈论的多智能体学习问题,在学习过程中,反馈信息可能会丢失,奖励是由一类广泛的游戏,称为变分稳定游戏。我们提出了一个简单的变种的经典在线梯度下降算法,称为重新加权在线梯度下降(ROGD),并表明,在变分稳定的游戏,如果每个代理采用ROGD,那么几乎肯定收敛到纳什均衡集是有保证的,即使当反馈损失是异步的,任意corrrelated代理之间。然后,我们扩展的框架,以处理未知的反馈损失概率,通过使用估计(从过去的数据构建)在其更换。最后,我们进一步扩展了框架,以适应异步损失和随机奖励,并建立多智能体ROGD学习仍然收敛到纳什均衡集在这样的设置。总之,这些结果有助于多代理在线学习的广泛景观显着放松所需的反馈信息,以实现理想的结果。
We consider a game-theoretical multi-agent learning problem where the feedback information can be lost during the learning process and rewards are given by a broad class of games known as variationally stable games. We propose a simple variant of the classical online gradient descent algorithm, called reweighted online gradient descent (ROGD) and show that in variationally stable games, if each agent adopts ROGD, then almost sure convergence to the set of Nash equilibria is guaranteed, even when the feedback loss is asynchronous and arbitrarily corrrelated among agents. We then extend the framework to deal with unknown feedback loss probabilities by using an estimator (constructed from past data) in its replacement. Finally, we further extend the framework to accomodate both asynchronous loss and stochastic rewards and establish that multi-agent ROGD learning still converges to the set of Nash equilibria in such settings. Together, these results contribute to the broad lanscape of multi-agent online learning by significantly relaxing the feedback information that is required to achieve desirable outcomes.