Approximating Discontinuous Nash Equilibrial Values of Two-Player General-Sum Differential Games

Approximating Discontinuous Nash Equilibrial Values of Two-Player General-Sum Differential Games
复制标题

DOI:
10.1109/icra48891.2023.10160219
复制
发表时间:
2022-07
期刊:
2023 IEEE International Conference on Robotics and Automation (ICRA)
影响因子:
--
通讯作者:
Lei Zhang;Mukesh Ghimire;Wenlong Zhang;Zhenni Xu;Yi Ren-
Lei Zhang;Mukesh Ghimire;Wenlong Zhang;Zhenni Xu;Yi Ren-
中科院分区:
其他
文献类型:
--
作者:
Lei Zhang;Mukesh Ghimire;Wenlong Zhang;Zhenni Xu;Yi Ren-

文献摘要

相似文献

寻找两人差分博弈的纳什均衡策略需要求解 Hamilton-Jacobi-Isaacs (HJI) PDE。自监督学习已被用来近似解决此类偏微分方程,同时避免维数灾难。然而,由于其采样性质,该方法无法学习不连续的 PDE 解决方案,导致当玩家奖励不连续时,机器人应用中所得控制器的安全性能较差。本文研究了该问题的两种潜在解决方案:一种利用监督纳什均衡和 HJI PDE 的混合方法,以及一种价值强化方法,其中通过逐渐强化奖励来解决一系列 HJI。我们使用分别具有 5D 和 9D 状态空间的两个车辆交互模拟研究中产生的泛化和安全性能来比较这些解决方案。结果表明,凭借信息丰富的监督(例如碰撞和接近碰撞演示)和自监督学习的低成本,混合方法在相同的计算预算下比监督、自监督和价值强化方法实现了更好的安全性能。在没有信息监督的情况下,价值强化无法在高维情况下泛化。最后,我们表明神经激活函数需要连续可微才能学习偏微分方程,并且其选择可以取决于具体情况。
Finding Nash equilibrial policies for two-player differential games requires solving Hamilton-Jacobi-Isaacs (HJI) PDEs. Self-supervised learning has been used to approximate solutions of such PDEs while circumventing the curse of dimensionality. However, this method fails to learn discontinuous PDE solutions due to its sampling nature, leading to poor safety performance of the resulting controllers in robotics applications when player rewards are discontinuous. This paper investigates two potential solutions to this problem: a hybrid method that leverages both supervised Nash equilibria and the HJI PDE, and a value-hardening method where a sequence of HJIs are solved with a gradually hardening reward. We compare these solutions using the resulting generalization and safety performance in two vehicle interaction simulation studies with 5D and 9D state spaces, respectively. Results show that with informative supervision (e.g., collision and near-collision demonstrations) and the low cost of self-supervised learning, the hybrid method achieves better safety performance than the supervised, self-supervised, and value hardening approaches on equal computational budget. Value hardening fails to generalize in the higher-dimensional case without informative supervision. Lastly, we show that the neural activation function needs to be continuously differentiable for learning PDEs and its choice can be case dependent.