Off-Policy Imitation Learning from Observations

Off-Policy Imitation Learning from Observations
复制标题

DOI:
--
复制
发表时间:
2021-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Zhuangdi Zhu;Kaixiang Lin;Bo Dai;Jiayu Zhou
Zhuangdi Zhu;Kaixiang Lin;Bo Dai;Jiayu Zhou
中科院分区:
其他
文献类型:
--
作者:
Zhuangdi Zhu;Kaixiang Lin;Bo Dai;Jiayu Zhou

文献摘要

被引文献

相似文献

从观察中学习(LfO)是一种实用的强化学习场景,许多应用程序可以通过重用不完整的资源而从中受益。与传统的模仿学习(IL)相比,LfO由于缺乏专家的行动指导而更具挑战性。在传统IL和LfO中,分布匹配都是其基础的核心。传统的分布匹配方法依赖于策略学习的非策略转换,样本成本高。对于样本效率,已经提出了一些非政策解决方案,然而,这些解决方案要么缺乏全面的理论依据,要么依赖于专家行动的指导。在这项工作中,我们提出了一种样本高效的LfO方法,该方法以原则性的方式实现了非策略优化。为了进一步加快学习过程,我们用一个逆动作模型来调节策略更新,从模式覆盖的角度帮助分布匹配。在具有挑战性的运动任务上的广泛实证结果表明,我们的方法在样本效率和渐近性能方面与最先进的方法相当。
Learning from Observations (LfO) is a practical reinforcement learning scenario from which many applications can benefit through the reuse of incomplete resources. Compared to conventional imitation learning (IL), LfO is more challenging because of the lack of expert action guidance. In both conventional IL and LfO, distribution matching is at the heart of their foundation. Traditional distribution matching approaches are sample-costly which depend on on-policy transitions for policy learning. Towards sample-efficiency, some off-policy solutions have been proposed, which, however, either lack comprehensive theoretical justifications or depend on the guidance of expert actions. In this work, we propose a sample-efficient LfO approach that enables off-policy optimization in a principled manner. To further accelerate the learning procedure, we regulate the policy update with an inverse action model, which assists distribution matching from the perspective of mode-covering. Extensive empirical results on challenging locomotion tasks indicate that our approach is comparable with state-of-the-art in terms of both sample-efficiency and asymptotic performance.