Stackelberg Actor-Critic: Game-Theoretic Reinforcement Learning Algorithms

Stackelberg Actor-Critic: Game-Theoretic Reinforcement Learning Algorithms
复制标题

DOI:
10.1609/aaai.v36i8.20908
复制
发表时间:
2021-09
期刊:
ArXiv
影响因子:
--
通讯作者:
Liyuan Zheng;Tanner Fiez;Zane Alumbaugh;Benjamin J. Chasnov;L. Ratliff
Liyuan Zheng;Tanner Fiez;Zane Alumbaugh;Benjamin J. Chasnov;L. Ratliff
中科院分区:
其他
文献类型:
--
作者:
Liyuan Zheng;Tanner Fiez;Zane Alumbaugh;Benjamin J. Chasnov;L. Ratliff

文献摘要

相似文献

在基于演员-评论家的强化学习算法中,演员和评论家之间的分层交互自然适合于博弈论的解释。我们采用这种观点和模型的演员和评论家的互动作为一个两个球员的一般和游戏的领导者,追随者的结构被称为Stackelberg游戏。鉴于这种抽象,我们提出了一个元框架Stackelberg演员批评家算法的领导者球员遵循其目标的总导数,而不是通常的个人梯度。从理论的角度来看,我们开发了一个政策梯度定理的精细更新,并提供了一个局部收敛保证Stackelberg演员批评算法的局部Stackelberg平衡。从经验的角度来看,我们通过简单的例子证明,我们研究的学习动态减轻循环和加速收敛相比,通常的梯度动态引起的演员-评论家配方的成本结构。最后,在OpenAI健身房环境中进行的大量实验表明,Stackelberg actor-critic算法的性能始终至少与标准actor-critic算法相当,并且通常显著优于标准actor-critic算法。
The hierarchical interaction between the actor and critic in actor-critic based reinforcement learning algorithms naturally lends itself to a game-theoretic interpretation. We adopt this viewpoint and model the actor and critic interaction as a two-player general-sum game with a leader-follower structure known as a Stackelberg game. Given this abstraction, we propose a meta-framework for Stackelberg actor-critic algorithms where the leader player follows the total derivative of its objective instead of the usual individual gradient. From a theoretical standpoint, we develop a policy gradient theorem for the refined update and provide a local convergence guarantee for the Stackelberg actor-critic algorithms to a local Stackelberg equilibrium. From an empirical standpoint, we demonstrate via simple examples that the learning dynamics we study mitigate cycling and accelerate convergence compared to the usual gradient dynamics given cost structures induced by actor-critic formulations. Finally, extensive experiments on OpenAI gym environments show that Stackelberg actor-critic algorithms always perform at least as well and often significantly outperform the standard actor-critic algorithm counterparts.