Efficient Iterative Linear-Quadratic Approximations for Nonlinear Multi-Player General-Sum Differential Games

Efficient Iterative Linear-Quadratic Approximations for Nonlinear Multi-Player General-Sum Differential Games
复制标题

非线性多人一般和微分博弈的高效迭代线性二次逼近

DOI:
--
复制
发表时间:
2019
期刊:
IEEE International Conference on Robotics and Automation
影响因子:
--
通讯作者:
C. Tomlin
C. Tomlin
中科院分区:
--
文献类型:
--
作者:
David Fridovich;Ellis Ratner;Lasse Peters;A. Dragan;C. Tomlin

文献摘要

被引文献

相似文献

机器人技术中的许多问题涉及多个决策主体。要在此类环境中高效运行,机器人必须考虑其决策对其他主体行为的影响。微分博弈为阐述这类多主体问题提供了一个富有表现力的理论框架。不幸的是,大多数数值解法在状态维度方面扩展性很差,很少用于实时应用。因此,通常会预测其他主体的未来决策,并解决由此产生的解耦(即单主体)最优控制问题。这种解耦忽略了问题潜在的交互性质;然而,对于广泛的最优控制问题类别,确实存在有效的解法。我们从一种这样的技术——迭代线性二次调节器(ILQR)中获得灵感,它解决具有线性动力学和二次成本的重复近似问题。类似地,我们提出的算法解决重复线性二次博弈。我们在多个具有各种初始条件的示例中对我们的算法进行实验基准测试,并表明所得策略呈现出复杂的交互行为。我们的结果表明我们的算法可靠收敛且能实时运行。在一个三主体、14状态的模拟交叉路口问题中,我们的算法最初在<0.25秒内收敛。在硬件防撞测试中,滚动时域调用在<50毫秒内收敛。
Many problems in robotics involve multiple decision making agents. To operate efficiently in such settings, a robot must reason about the impact of its decisions on the behavior of other agents. Differential games offer an expressive theoretical framework for formulating these types of multi-agent problems. Unfortunately, most numerical solution techniques scale poorly with state dimension and are rarely used in real-time applications. For this reason, it is common to predict the future decisions of other agents and solve the resulting decoupled, i.e., single-agent, optimal control problem. This decoupling neglects the underlying interactive nature of the problem; however, efficient solution techniques do exist for broad classes of optimal control problems. We take inspiration from one such technique, the iterative linear-quadratic regulator (ILQR), which solves repeated approximations with linear dynamics and quadratic costs. Similarly, our proposed algorithm solves repeated linear-quadratic games. We experimentally benchmark our algorithm in several examples with a variety of initial conditions and show that the resulting strategies exhibit complex interactive behavior. Our results indicate that our algorithm converges reliably and runs in real-time. In a three-player, 14-state simulated intersection problem, our algorithm initially converges in < 0.25 s. Receding horizon invocations converge in < 50 ms in a hardware collision-avoidance test.