Maximum-Entropy Multi-Agent Dynamic Games: Forward and Inverse Solutions

Maximum-Entropy Multi-Agent Dynamic Games: Forward and Inverse Solutions
复制标题

DOI:
10.1109/tro.2022.3232300
复制
发表时间:
2021-10
影响因子:
7.8
通讯作者:
Negar Mehr;Mingyu Wang;Maulik Bhatt;M. Schwager
Negar Mehr;Mingyu Wang;Maulik Bhatt;M. Schwager
中科院分区:
计算机科学1区
文献类型:
--
作者:
Negar Mehr;Mingyu Wang;Maulik Bhatt;M. Schwager

文献摘要

相似文献

在这篇文章中,我们研究了具有连续状态和动作空间的动态博弈场景中多个随机主体相互作用的问题。我们定义了有限理性主体的随机纳什均衡的一个新概念,我们称之为熵成本均衡。我们证明了ECA是对单个智能体具有最大熵最优性的多个智能体的自然扩展。我们解决了多智能体欧洲经委会博弈的“正向”和“反向”问题。对于前向问题,我们给出了一种Riccati算法来计算智能体的闭式ECA反馈策略,该策略在线性-二次-高斯情形下是精确的。我们给出了一个迭代变量来寻找非线性情形下的欧洲经委会局部反馈策略。对于反问题,我们给出了一个算法来推断多个相互作用的代理的成本函数,这些代理给出了噪声、有界合理的输入和状态轨迹示例。在一个模拟的多智能体碰撞避免场景中,并使用来自交互流量数据集的数据,我们的算法的有效性得到了验证。在这两种情况下,我们证明,与标准的逆最优控制方法相比,通过使用我们的算法来考虑代理人的博弈论相互作用,可以学习到更准确的代理人成本模型。
In this article, we study the problem of multiple stochastic agents interacting in a dynamic game scenario with continuous state and action spaces. We define a new notion of stochastic Nash equilibrium for boundedly rational agents, which we call the entropic cost equilibrium (ECE). We show that ECE is a natural extension to multiple agents of maximum entropy optimality for a single agent. We solve both the “forward” and “inverse” problems for the multi-agent ECE game. For the forward problem, we provide a Riccati algorithm to compute closed-form ECE feedback policies for the agents, which are exact in the linear-quadratic-gaussian case. We give an iterative variant to find locally ECE feedback policies for the nonlinear case. For the inverse problem, we present an algorithm to infer the cost functions of the multiple interacting agents given noisy, boundedly rational input and state trajectory examples from agents acting in an ECE. The effectiveness of our algorithms is demonstrated in a simulated multi-agent collision avoidance scenario, and with data from the INTERACTION traffic dataset. In both cases, we show that, by taking into account the agents' game theoretic interactions using our algorithm, a more accurate model of agents' costs can be learned, compared with standard inverse optimal control methods.