DEALIO: Data-Efficient Adversarial Learning for Imitation from Observation

DEALIO: Data-Efficient Adversarial Learning for Imitation from Observation
复制标题

DOI:
10.1109/iros51168.2021.9636169
复制
发表时间:
2021-03
期刊:
2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
F. Torabi;Garrett Warnell;P. Stone
F. Torabi;Garrett Warnell;P. Stone
中科院分区:
其他
文献类型:
--
作者:
F. Torabi;Garrett Warnell;P. Stone

文献摘要

被引文献

相似文献

在从观察中模仿学习(IfO)中,学习代理只使用对演示行为的观察来模仿演示代理,而不访问演示者生成的控制信号。最近基于对抗性模仿学习的方法在IfO问题上取得了最先进的性能,但由于依赖于数据效率低下、无模型的强化学习算法,它们通常具有高样本复杂度。这个问题使得它们在现实环境中部署不切实际,在现实环境中,收集样本可能会在时间,能源和风险方面产生高成本。在这项工作中,我们假设我们可以将基于模型的强化学习的思想与IfO的对抗方法相结合,以便在不牺牲性能的情况下提高这些方法的数据效率。具体来说,我们考虑时变线性高斯政策,并提出了一种方法,集成了线性二次调节器与路径积分政策改进到现有的对抗IfO框架。其结果是一个更有效的数据IfO算法具有更好的性能,我们在四个模拟领域的经验表明:使用少得多的与环境的相互作用,所提出的方法表现出类似或更好的性能比现有技术。
In imitation learning from observation (IfO), a learning agent seeks to imitate a demonstrating agent using only observations of the demonstrated behavior without access to the control signals generated by the demonstrator. Recent methods based on adversarial imitation learning have led to state-of-the-art performance on IfO problems, but they typically suffer from high sample complexity due to a reliance on data-inefficient, model-free reinforcement learning algorithms. This issue makes them impractical to deploy in real-world settings, where gathering samples can incur high costs in terms of time, energy, and risk. In this work, we hypothesize that we can incorporate ideas from model-based reinforcement learning with adversarial methods for IfO in order to increase the data efficiency of these methods without sacrificing performance. Specifically, we consider time-varying linear Gaussian policies, and propose a method that integrates the linear-quadratic regulator with path integral policy improvement into an existing adversarial IfO framework. The result is a more data-efficient IfO algorithm with better performance, which we show empirically in four simulation domains: using far fewer interactions with the environment, the proposed method exhibits similar or better performance than the existing technique.