A Ranking Game for Imitation Learning

A Ranking Game for Imitation Learning
复制标题

DOI:
--
复制
发表时间:
2022-02
期刊:
ArXiv
影响因子:
--
通讯作者:
Harshit S. Sikchi;Akanksha Saran;Wonjoon Goo;S. Niekum
Harshit S. Sikchi;Akanksha Saran;Wonjoon Goo;S. Niekum
中科院分区:
其他
文献类型:
--
作者:
Harshit S. Sikchi;Akanksha Saran;Wonjoon Goo;S. Niekum

文献摘要

被引文献

相似文献

我们提出了一个新的模仿学习框架——将模仿视为政策和奖励之间基于排名的两人游戏。在这个游戏中,奖励代理学习满足行为之间的成对绩效排名,而策略代理学习最大化该奖励。在模仿学习中,接近最优的专家数据可能很难获得,即使在无限数据的限制下也不能像偏好那样暗示对轨迹的总排序。另一方面,仅从偏好中学习具有挑战性,因为需要大量偏好来推断高维奖励函数,尽管偏好数据通常比专家演示更容易收集。经典的逆强化学习(IRL)公式从专家演示中学习,但没有提供结合离线偏好学习的机制,反之亦然。我们用新颖的排名损失实例化了所提出的排名游戏框架,给出了一种可以同时从专家演示和偏好中学习的算法,从而获得了两种模式的优点。我们的实验表明,所提出的方法实现了最先进的样本效率,并且可以解决以前在观察学习(LfO)设置中无法解决的任务。项目视频和代码可以在 https://hari-sikchi.github.io/rank-game/ 找到
We propose a new framework for imitation learning -- treating imitation as a two-player ranking-based game between a policy and a reward. In this game, the reward agent learns to satisfy pairwise performance rankings between behaviors, while the policy agent learns to maximize this reward. In imitation learning, near-optimal expert data can be difficult to obtain, and even in the limit of infinite data cannot imply a total ordering over trajectories as preferences can. On the other hand, learning from preferences alone is challenging as a large number of preferences are required to infer a high-dimensional reward function, though preference data is typically much easier to collect than expert demonstrations. The classical inverse reinforcement learning (IRL) formulation learns from expert demonstrations but provides no mechanism to incorporate learning from offline preferences and vice versa. We instantiate the proposed ranking-game framework with a novel ranking loss giving an algorithm that can simultaneously learn from expert demonstrations and preferences, gaining the advantages of both modalities. Our experiments show that the proposed method achieves state-of-the-art sample efficiency and can solve previously unsolvable tasks in the Learning from Observation (LfO) setting. Project video and code can be found at https://hari-sikchi.github.io/rank-game/