Independent Policy Gradient Methods for Competitive Reinforcement Learning

Independent Policy Gradient Methods for Competitive Reinforcement Learning
复制标题

DOI:
--
复制
发表时间:
2021-01
期刊:
ArXiv
影响因子:
--
通讯作者:
C. Daskalakis;Dylan J. Foster;Noah Golowich
C. Daskalakis;Dylan J. Foster;Noah Golowich
中科院分区:
其他
文献类型:
--
作者:
C. Daskalakis;Dylan J. Foster;Noah Golowich

文献摘要

被引文献

相似文献

在具有两个智能体(即零和随机博弈)的竞争强化学习环境中,我们为独立学习算法获得了全局的、非渐近收敛保证。我们考虑一种情节式设置,在每个情节中,每个参与者独立地选择一个策略,并且仅观察他们自己的行动和奖励以及状态。我们表明,如果两个参与者协同运行策略梯度方法,只要他们的学习率遵循双时间尺度规则(这是必要的),他们的策略将收敛到博弈的最小 - 最大均衡。据我们所知,这是竞争强化学习中独立策略梯度方法的第一个有限样本收敛结果;先前的工作主要集中在用于均衡计算的集中式、协调式过程。
We obtain global, non-asymptotic convergence guarantees for independent learning algorithms in competitive reinforcement learning settings with two agents (i.e., zero-sum stochastic games). We consider an episodic setting where in each episode, each player independently selects a policy and observes only their own actions and rewards, along with the state. We show that if both players run policy gradient methods in tandem, their policies will converge to a min-max equilibrium of the game, as long as their learning rates follow a two-timescale rule (which is necessary). To the best of our knowledge, this constitutes the first finite-sample convergence result for independent policy gradient methods in competitive RL; prior work has largely focused on centralized, coordinated procedures for equilibrium computation.