On Computation and Generalization of Generative Adversarial Imitation Learning

On Computation and Generalization of Generative Adversarial Imitation Learning
复制标题

DOI:
--
复制
发表时间:
2020-01
期刊:
ArXiv
影响因子:
--
通讯作者:
Minshuo Chen;Yizhou Wang;Tianyi Liu;Zhuoran Yang;Xingguo Li;Zhaoran Wang;T. Zhao
Minshuo Chen;Yizhou Wang;Tianyi Liu;Zhuoran Yang;Xingguo Li;Zhaoran Wang;T. Zhao
中科院分区:
其他
文献类型:
--
作者:
Minshuo Chen;Yizhou Wang;Tianyi Liu;Zhuoran Yang;Xingguo Li;Zhaoran Wang;T. Zhao

文献摘要

被引文献

相似文献

生成对抗模仿学习(GAIL)是一种强大而实用的学习顺序决策策略的方法。与强化学习(RL)不同,GAIL利用专家的演示数据(例如,人类),并学习未知环境的策略和奖励函数。尽管取得了重大的经验进展,GAIL背后的理论在很大程度上仍然是未知的。主要的困难来自潜在的时间依赖性的示范数据和极小极大计算公式GAIL没有凹凸结构。为了弥合理论与实践之间的差距,本文研究了GAIL的理论特性。具体而言,我们显示:(1)对于一般报酬参数化的GAIL,只要适当控制报酬函数的类,就可以保证推广性;(2)对于报酬参数化为再生核函数的GAIL,可以用随机一阶优化算法有效地求解,该算法可以达到次线性收敛到平稳解.据我们所知,这些是第一个结果的统计和计算保证的模仿学习与奖励/政策函数近似。数值实验支持我们的分析。
Generative Adversarial Imitation Learning (GAIL) is a powerful and practical approach for learning sequential decision-making policies. Different from Reinforcement Learning (RL), GAIL takes advantage of demonstration data by experts (e.g., human), and learns both the policy and reward function of the unknown environment. Despite the significant empirical progresses, the theory behind GAIL is still largely unknown. The major difficulty comes from the underlying temporal dependency of the demonstration data and the minimax computational formulation of GAIL without convex-concave structure. To bridge such a gap between theory and practice, this paper investigates the theoretical properties of GAIL. Specifically, we show: (1) For GAIL with general reward parameterization, the generalization can be guaranteed as long as the class of the reward functions is properly controlled; (2) For GAIL, where the reward is parameterized as a reproducing kernel function, GAIL can be efficiently solved by stochastic first order optimization algorithms, which attain sublinear convergence to a stationary solution. To the best of our knowledge, these are the first results on statistical and computational guarantees of imitation learning with reward/policy function ap- proximation. Numerical experiments are provided to support our analysis.