Model-Free Mean-Field Reinforcement Learning: Mean-Field MDP and Mean-Field Q-Learning

Model-Free Mean-Field Reinforcement Learning: Mean-Field MDP and Mean-Field Q-Learning
复制标题

DOI:
10.1214/23-aap1949
复制
发表时间:
2019-10
期刊:
ArXiv
影响因子:
--
通讯作者:
R. Carmona;M. Laurière;Zongjun Tan
R. Carmona;M. Laurière;Zongjun Tan
中科院分区:
其他
文献类型:
--
作者:
R. Carmona;M. Laurière;Zongjun Tan

文献摘要

被引文献

相似文献

针对平均场控制(MFC)问题,提出了一个通用的强化学习框架。例如,当多个智能体的数量很大时,就会出现协作多智能体控制问题的极限等问题。渐近问题可以归结为一个非线性动力学的最优控制问题。这也可以被视为马尔可夫决策过程(MDP),但与通常的RL设置的关键区别在于,动态和奖励现在取决于状态的概率分布本身。或者,它可以被重塑为度量的Wasserstein空间上的MDP。在这项工作中,我们引入了基于平均场水平的状态-作用值函数的通用无模型算法,并证明了一个典型的Q-学习方法的收敛。然后,我们实现了参与者-批评者方法,并报告了两个原型问题的数值结果:一个是由网络安全应用驱动的有限空间模型,另一个是由群运动应用驱动的连续空间模型。
We develop a general reinforcement learning framework for mean field control (MFC) problems. Such problems arise for instance as the limit of collaborative multi-agent control problems when the number of agents is very large. The asymptotic problem can be phrased as the optimal control of a non-linear dynamics. This can also be viewed as a Markov decision process (MDP) but the key difference with the usual RL setup is that the dynamics and the reward now depend on the state's probability distribution itself. Alternatively, it can be recast as a MDP on the Wasserstein space of measures. In this work, we introduce generic model-free algorithms based on the state-action value function at the mean field level and we prove convergence for a prototypical Q-learning method. We then implement an actor-critic method and report numerical results on two archetypal problems: a finite space model motivated by a cyber security application and a continuous space model motivated by an application to swarm motion.