When Will Generative Adversarial Imitation Learning Algorithms Attain Global Convergence

When Will Generative Adversarial Imitation Learning Algorithms Attain Global Convergence
复制标题

DOI:
--
复制
发表时间:
2020-06
期刊:
--
影响因子:
--
通讯作者:
Ziwei Guan;Tengyu Xu;Yingbin Liang
Ziwei Guan;Tengyu Xu;Yingbin Liang
中科院分区:
其他
文献类型:
--
作者:
Ziwei Guan;Tengyu Xu;Yingbin Liang

文献摘要

相似文献

生成对抗模仿学习(GAIL)是一种流行的反向强化学习方法,用于从专家轨迹中联合优化策略和奖励。关于GAIL的一个主要问题是将某种策略梯度算法应用于GAIL是否获得全局最小值(即,产生专家策略),对此现有的理解非常有限。这样的全局收敛性仅在线性(或线性型)MDP和线性(或可线性化)报酬下得到证明。在本文中,我们研究GAIL下一般MDP和非线性奖励函数类(只要目标函数是强凹的奖励参数)。我们的全局收敛性与广泛的常用的政策梯度算法,所有这些都是在交替的方式与随机梯度上升的奖励更新,包括预计的政策梯度(PPG)-GAIL,弗兰克-沃尔夫政策梯度(FWPG)-GAIL,信赖域政策优化(TRPO)-GAIL和自然政策梯度(NPG)-GAIL。这是第一次对GAIL算法的全局收敛性进行系统的理论研究。
Generative adversarial imitation learning (GAIL) is a popular inverse reinforcement learning approach for jointly optimizing policy and reward from expert trajectories. A primary question about GAIL is whether applying a certain policy gradient algorithm to GAIL attains a global minimizer (i.e., yields the expert policy), for which existing understanding is very limited. Such global convergence has been shown only for the linear (or linear-type) MDP and linear (or linearizable) reward. In this paper, we study GAIL under general MDP and for nonlinear reward function classes (as long as the objective function is strongly concave with respect to the reward parameter). We characterize the global convergence with a sublinear rate for a broad range of commonly used policy gradient algorithms, all of which are implemented in an alternating manner with stochastic gradient ascent for reward update, including projected policy gradient (PPG)-GAIL, Frank-Wolfe policy gradient (FWPG)-GAIL, trust region policy optimization (TRPO)-GAIL and natural policy gradient (NPG)-GAIL. This is the first systematic theoretical study of GAIL for global convergence.