Accelerating Imitation Learning with Predictive Models

Accelerating Imitation Learning with Predictive Models
复制标题

DOI:
--
复制
发表时间:
2018-06
期刊:
--
影响因子:
--
通讯作者:
Ching-An Cheng;Xinyan Yan;Evangelos A. Theodorou;Byron Boots
Ching-An Cheng;Xinyan Yan;Evangelos A. Theodorou;Byron Boots
中科院分区:
其他
文献类型:
--
作者:
Ching-An Cheng;Xinyan Yan;Evangelos A. Theodorou;Byron Boots

文献摘要

相似文献

样本效率对于解决现实世界的强化学习问题至关重要,因为在现实世界中,代理与环境的交互可能是昂贵的。事实证明,从专家建议中模仿学习是一种有效的策略,可以减少训练策略所需的交互次数。在线模仿学习将策略评估和策略优化交织在一起,是一种特别有效的技术,具有可证明的性能保证。在这项工作中,我们寻求进一步加快在线模仿学习的收敛速度,从而使其更具样本效率。我们提出了两个基于模型的算法的启发跟随领导者(FTL)与预测:MoBIL-VI的基础上解决变分不等式和MoBIL-Prox的基础上随机一阶更新。这两种方法利用模型来预测未来的梯度,以加速策略学习。当模型预言机在线学习时,这些算法可以证明将已知的最佳收敛速度加快一个数量级。我们的算法可以被看作是一个推广的随机逼近(Juditsky等人,2011年),并承认一个简单的建设性的超光速风格的性能分析。
Sample efficiency is critical in solving real-world reinforcement learning problems, where agent-environment interactions can be costly. Imitation learning from expert advice has proved to be an effective strategy for reducing the number of interactions required to train a policy. Online imitation learning, which interleaves policy evaluation and policy optimization, is a particularly effective technique with provable performance guarantees. In this work, we seek to further accelerate the convergence rate of online imitation learning, thereby making it more sample efficient. We propose two model-based algorithms inspired by Follow-the-Leader (FTL) with prediction: MoBIL-VI based on solving variational inequalities and MoBIL-Prox based on stochastic first-order updates. These two methods leverage a model to predict future gradients to speed up policy learning. When the model oracle is learned online, these algorithms can provably accelerate the best known convergence rate up to an order. Our algorithms can be viewed as a generalization of stochastic Mirror-Prox (Juditsky et al., 2011), and admit a simple constructive FTL-style analysis of performance.