Unified reinforcement Q-learning for mean field game and control problems

Unified reinforcement Q-learning for mean field game and control problems
复制标题

DOI:
10.1007/s00498-021-00310-1
复制
发表时间:
2020-06
期刊:
Mathematics of Control, Signals, and Systems
影响因子:
--
通讯作者:
Andrea Angiuli;J. Fouque;M. Laurière
Andrea Angiuli;J. Fouque;M. Laurière
中科院分区:
其他
文献类型:
--
作者:
Andrea Angiuli;J. Fouque;M. Laurière

文献摘要

相似文献

提出了一种求解无限时域渐近平均场博弈和平均场控制问题的强化学习算法。我们的方法可以被描述为一个统一的双时标平均场Q学习:Thesamealgorithm可以学习MFG或MFC的解决方案,通过简单地调整两个学习参数的比例。该算法是在离散的时间和空间,代理不仅提供了一个行动的环境,但也分布的状态,以考虑到平均场的功能的问题。重要的是,我们假设代理不能观察到人口的分布,需要以无模型的方式估计它。渐近MFG和MFC问题也提出了在连续的时间和空间,并与经典(非渐近或平稳)MFG和MFC问题。它们导致显式的解决方案,在线性二次(LQ)的情况下,作为基准,我们的算法的结果。
We present a Reinforcement Learning (RL) algorithm to solve infinite horizon asymptotic Mean Field Game (MFG) and Mean Field Control (MFC) problems. Our approach can be described as a unified two-timescale Mean Field Q-learning: Thesamealgorithm can learn either the MFG or the MFC solution by simply tuning the ratio of two learning parameters. The algorithm is in discrete time and space where the agent not only provides an action to the environment but also a distribution of the state in order to take into account the mean field feature of the problem. Importantly, we assume that the agent cannot observe the population’s distribution and needs to estimate it in a model-free manner. The asymptotic MFG and MFC problems are also presented in continuous time and space, and compared with classical (non-asymptotic or stationary) MFG and MFC problems. They lead to explicit solutions in the linear-quadratic (LQ) case that are used as benchmarks for the results of our algorithm.