Game Theory for Autonomy: From Min-Max Optimization to Equilibrium and Bounded Rationality Learning

Game Theory for Autonomy: From Min-Max Optimization to Equilibrium and Bounded Rationality Learning
复制标题

DOI:
10.23919/acc55779.2023.10156432
复制
发表时间:
2023-05
期刊:
2023 American Control Conference (ACC)
影响因子:
--
通讯作者:
K. Vamvoudakis;Filippos Fotiadis;J. Hespanha;Raphael Chinchilla;Guosong Yang;Mushuang Liu;J. Shamma;Lacra Pavel
K. Vamvoudakis;Filippos Fotiadis;J. Hespanha;Raphael Chinchilla;Guosong Yang;Mushuang Liu;J. Shamma;Lacra Pavel
中科院分区:
其他
文献类型:
--
作者:
K. Vamvoudakis;Filippos Fotiadis;J. Hespanha;Raphael Chinchilla;Guosong Yang;Mushuang Liu;J. Shamma;Lacra Pavel

文献摘要

被引文献

相似文献

一般来说,在非合作博弈中找到纳什均衡是一项极具挑战性的任务。这是由多种因素造成的,包括但不限于游戏的成本函数是非凸/非凹,游戏玩家之间的信息有限,甚至是计算复杂性的问题。本教程从这一残酷的现实中汲取动力,并提供了使用优化和基于学习的技术在非理想环境中近似纳什均衡或最小-最大均衡的方法。然而,本教程承认,这些技术可能并不总是收敛,而是导致振荡甚至混乱。在这方面,提供了被动和耗散理论的工具,可以解释这些不同的行为。最后,本教程强调,寻找均衡政策是徒劳的,这比人们通常认为的要频繁得多;相反,有限理性和非均衡策略可以更现实地使用,因为一些参与者的学习不完美或相对幼稚——“有限理性”。在自动驾驶系统的背景下,这些游戏的有效性得到了证明,在自动驾驶系统中,它们明确表明可以保证车辆安全。
Finding Nash equilibria in non-cooperative games can be, in general, an exceptionally challenging task. This is owed to various factors, including but not limited to the cost functions of the game being nonconvex/nonconcave, the players of the game having limited information about one another, or even due to issues of computational complexity. The present tutorial draws motivation from this harsh reality and provides methods to approximate Nash or min-max equilibria in non-ideal settings using both optimization- and learning-based techniques. The tutorial acknowledges, however, that such techniques may not always converge, but instead lead to oscillations or even chaos. In that respect, tools from passivity and dissipativity theory are provided, which can offer explanations about these divergent behaviors. Finally, the tutorial highlights that, more frequently than often thought, the search for equilibrium policies is simply vain; instead, bounded rationality and non-equilibrium policies can be more realistic to employ owing to some players’ learning imperfectly or being relatively naive – "bounded rational." The efficacy of such plays is demonstrated in the context of autonomous driving systems, where it is explicitly shown that they can guarantee vehicle safety.