Fast Policy Learning through Imitation and Reinforcement

Fast Policy Learning through Imitation and Reinforcement
复制标题

DOI:
--
复制
发表时间:
2018-05
期刊:
ArXiv
影响因子:
--
通讯作者:
Ching-An Cheng;Xinyan Yan;Nolan Wagener;Byron Boots
Ching-An Cheng;Xinyan Yan;Nolan Wagener;Byron Boots
中科院分区:
其他
文献类型:
--
作者:
Ching-An Cheng;Xinyan Yan;Nolan Wagener;Byron Boots

文献摘要

被引文献

相似文献

模仿学习(IL)由一系列利用专家演示来快速学习策略的工具组成。然而,如果专家并非最优,与强化学习(RL)相比,模仿学习可能产生性能较差的策略。在本文中,我们旨在提供一种结合强化学习和模仿学习最佳方面的算法。我们通过在一个通用的镜像下降框架中构建几种流行的强化学习和模仿学习算法来实现这一点,表明这些算法可被视为一种单一方法的变体。然后我们提出洛基(LOKI),这是一种策略学习策略,它在切换到策略梯度强化学习方法之前,先执行少量但随机数量的模仿学习迭代。我们表明,如果切换时间适当随机化,洛基能够学会超越非最优专家,并且比从头开始运行策略梯度收敛得更快。最后,我们在几个模拟环境中通过实验评估洛基的性能。
Imitation learning (IL) consists of a set of tools that leverage expert demonstrations to quickly learn policies. However, if the expert is suboptimal, IL can yield policies with inferior performance compared to reinforcement learning (RL). In this paper, we aim to provide an algorithm that combines the best aspects of RL and IL. We accomplish this by formulating several popular RL and IL algorithms in a common mirror descent framework, showing that these algorithms can be viewed as a variation on a single approach. We then propose LOKI, a strategy for policy learning that first performs a small but random number of IL iterations before switching to a policy gradient RL method. We show that if the switching time is properly randomized, LOKI can learn to outperform a suboptimal expert and converge faster than running policy gradient from scratch. Finally, we evaluate the performance of LOKI experimentally in several simulated environments.