Large population stochastic dynamic games: closed-loop McKean-Vlasov systems and the Nash certainty equivalence principle

Large population stochastic dynamic games: closed-loop McKean-Vlasov systems and the Nash certainty equivalence principle
复制标题

DOI:
10.4310/cis.2006.v6.n3.a5
复制
发表时间:
2006-10
期刊:
Commun. Inf. Syst.
影响因子:
--
通讯作者:
P. Caines;Minyi Huang;R. Malhamé
P. Caines;Minyi Huang;R. Malhamé
中科院分区:
其他
文献类型:
--
作者:
P. Caines;Minyi Huang;R. Malhamé

文献摘要

被引文献

相似文献

摘要。我们考虑大种群条件下的随机动态博弈,其中多类智能体通过其个体动态和成本弱耦合。我们通过所谓的纳什确定性等价(NCE)原则来解决这个大群体博弈问题,该原则导致分散控制综合。本文提出的McKean-Vlasov NCE方法与大粒子系统的统计物理有密切的联系:两者都在微观层面上确定个体主体(或粒子)与宏观层面上个体(或粒子)质量之间的一致性关系。整个博弈被分解为(i)一个最优控制问题,其Hamilton-Jacobi-Bellman (HJB)方程决定了每个个体的最优控制,并涉及与质量效应相对应的度量,以及(ii)一系列McKean-Vlasov (M-V)方程,也依赖于该度量。我们将NCE原理定义为结果方案是一致(或可溶)的性质,即规定的控制律产生产生质量效应测量的样本路径。通过构造,整个闭环行为使得每个代理的行为在博弈纳什意义上相对于所有其他代理是最优的。
Abstract. We consider stochastic dynamic games in large population conditions where multiclass agents are weakly coupled via their individual dynamics and costs. We approach this large population game problem by the so-called Nash Certainty Equivalence (NCE) Principle which leads to a decentralized control synthesis. The McKean-Vlasov NCE method presented in this paper has a close connection with the statistical physics of large particle systems: both identify a consistency relationship between the individual agent (or particle) at the microscopic level and the mass of individuals (or particles) at the macroscopic level. The overall game is decomposed into (i) an optimal control problem whose Hamilton-Jacobi-Bellman (HJB) equation determines the optimal control for each individual and which involves a measure corresponding to the mass effect, and (ii) a family of McKean-Vlasov (M-V) equations which also depend upon this measure. We designate the NCE Principle as the property that the resulting scheme is consistent (or soluble), i.e. the prescribed control laws produce sample paths which produce the mass effect measure. By construction, the overall closed-loop behaviour is such that each agent’s behaviour is optimal with respect to all other agents in the game theoretic Nash sense.