Multi-Agent Reinforcement Learning in Cournot Games

Multi-Agent Reinforcement Learning in Cournot Games
复制标题

DOI:
10.1109/cdc42340.2020.9304089
复制
发表时间:
2020-09
期刊:
2020 59th IEEE Conference on Decision and Control (CDC)
影响因子:
--
通讯作者:
Yuanyuan Shi;Baosen Zhang
Yuanyuan Shi;Baosen Zhang
中科院分区:
其他
文献类型:
--
作者:
Yuanyuan Shi;Baosen Zhang

文献摘要

被引文献

相似文献

在这项工作中,我们研究了信息反馈有限的连续行动古诺博弈中策略主体的相互作用。古诺博弈是许多社会经济系统的基本市场模型,在该系统中,主体在不完全了解系统或彼此的情况下学习和竞争。我们考虑凹古诺博弈中策略梯度算法的动态性,策略梯度算法是一种广泛采用的连续控制强化学习算法。我们证明当价格函数为线性或代理数量为 2 时,政策梯度动态收敛于纳什均衡。这是关于具有不属于无悔类别的连续动作空间的学习算法的收敛特性的第一个结果(据我们所知)。
In this work, we study the interaction of strategic agents in continuous action Cournot games with limited information feedback. Cournot game is the essential market model for many socio-economic systems where agents learn and compete without the full knowledge of the system or each other. We consider the dynamics of the policy gradient algorithm, which is a widely adopted continuous control reinforcement learning algorithm, in concave Cournot games. We prove the convergence of policy gradient dynamics to the Nash equilibrium when the price function is linear or the number of agents is two. This is the first result (to the best of our knowledge) on the convergence property of learning algorithms with continuous action spaces that do not fall in the no-regret class.