Extrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations

Extrapolating Beyond Suboptimal Demonstrations via Inverse Reinforcement Learning from Observations
复制标题

DOI:
--
复制
发表时间:
2019-04
期刊:
--
影响因子:
--
通讯作者:
Daniel S. Brown;Wonjoon Goo;P. Nagarajan;S. Niekum
Daniel S. Brown;Wonjoon Goo;P. Nagarajan;S. Niekum
中科院分区:
其他
文献类型:
--
作者:
Daniel S. Brown;Wonjoon Goo;P. Nagarajan;S. Niekum

文献摘要

被引文献

相似文献

现有逆强化学习(IRL)方法的一个关键缺陷是它们无法显著优于演示器。这是因为IRL通常寻求一个奖励函数,使演示者看起来接近最优,而不是推断演示者在实践中可能执行得很差的潜在意图。在本文中,我们引入了一种新的从观察中学习的奖励算法,轨迹排名奖励外推(T-REX),该算法在一组(近似)排名的演示之外进行外推,以便从一组可能较差的演示中推断出高质量的奖励函数。当与深度强化学习相结合时,T-REX在多个Atari和MuJoCo基准任务上的表现优于最先进的模仿学习和IRL方法,并且通常达到最佳演示性能的两倍以上。我们还证明了T-REX对噪声排序的鲁棒性,并且可以通过简单地观察学习者随着时间的推移在任务中有噪声的改进来准确地推断意图。
A critical flaw of existing inverse reinforcement learning (IRL) methods is their inability to significantly outperform the demonstrator. This is because IRL typically seeks a reward function that makes the demonstrator appear near-optimal, rather than inferring the underlying intentions of the demonstrator that may have been poorly executed in practice. In this paper, we introduce a novel reward-learning-from-observation algorithm, Trajectory-ranked Reward EXtrapolation (T-REX), that extrapolates beyond a set of (approximately) ranked demonstrations in order to infer high-quality reward functions from a set of potentially poor demonstrations. When combined with deep reinforcement learning, T-REX outperforms state-of-the-art imitation learning and IRL methods on multiple Atari and MuJoCo benchmark tasks and achieves performance that is often more than twice the performance of the best demonstration. We also demonstrate that T-REX is robust to ranking noise and can accurately extrapolate intention by simply watching a learner noisily improve at a task over time.