Better-than-Demonstrator Imitation Learning via Automatically-Ranked Demonstrations

Better-than-Demonstrator Imitation Learning via Automatically-Ranked Demonstrations
复制标题

DOI:
--
复制
发表时间:
2019-07
期刊:
--
影响因子:
--
通讯作者:
Daniel S. Brown;Wonjoon Goo;S. Niekum
Daniel S. Brown;Wonjoon Goo;S. Niekum
中科院分区:
其他
文献类型:
--
作者:
Daniel S. Brown;Wonjoon Goo;S. Niekum

文献摘要

被引文献

相似文献

模仿学习的性能通常是由演示者的性能上限。虽然最近的实证结果表明,排名示范允许更好的性能比演示,演示的偏好可能很难获得,理论上知之甚少,当这种方法可以预期成功地外推超出演示的性能。为了解决这些问题,我们首先为优于演示者的模仿学习提供了一个充分条件,并提供了理论结果,说明为什么在执行逆强化学习时,对演示的偏好可以更好地减少奖励函数的模糊性。在此理论的基础上,我们引入了基于干扰的奖励外推(D-REX),这是一种基于排名的模仿学习方法,它将噪声注入通过行为克隆学习的策略中,以自动生成排名演示。这些排名的演示用于有效地学习奖励函数,然后可以使用强化学习进行优化。我们在模拟机器人和Atari模仿学习基准上实证验证了我们的方法,并表明D-REX优于标准模仿学习方法,并且可以显着超过演示器的性能。D-REX是第一种模仿学习方法,可以在没有额外的辅助信息或监督(如奖励或人类偏好)的情况下实现对演示者表现的显著外推。通过自动生成排名,我们证明了基于偏好的逆强化学习可以应用于传统的模仿学习环境,其中只有未标记的演示可用。
The performance of imitation learning is typically upper-bounded by the performance of the demonstrator. While recent empirical results demonstrate that ranked demonstrations allow for better-than-demonstrator performance, preferences over demonstrations may be difficult to obtain, and little is known theoretically about when such methods can be expected to successfully extrapolate beyond the performance of the demonstrator. To address these issues, we first contribute a sufficient condition for better-than-demonstrator imitation learning and provide theoretical results showing why preferences over demonstrations can better reduce reward function ambiguity when performing inverse reinforcement learning. Building on this theory, we introduce Disturbance-based Reward Extrapolation (D-REX), a ranking-based imitation learning method that injects noise into a policy learned through behavioral cloning to automatically generate ranked demonstrations. These ranked demonstrations are used to efficiently learn a reward function that can then be optimized using reinforcement learning. We empirically validate our approach on simulated robot and Atari imitation learning benchmarks and show that D-REX outperforms standard imitation learning approaches and can significantly surpass the performance of the demonstrator. D-REX is the first imitation learning approach to achieve significant extrapolation beyond the demonstrator's performance without additional side-information or supervision, such as rewards or human preferences. By generating rankings automatically, we show that preference-based inverse reinforcement learning can be applied in traditional imitation learning settings where only unlabeled demonstrations are available.