Learning Reward Functions by Integrating Human Demonstrations and Preferences

Learning Reward Functions by Integrating Human Demonstrations and Preferences
复制标题

DOI:
10.15607/rss.2019.xv.023
复制
发表时间:
2019-06
期刊:
ArXiv
影响因子:
--
通讯作者:
Malayandi Palan;Nicholas C. Landolfi;Gleb Shevchuk;Dorsa Sadigh
Malayandi Palan;Nicholas C. Landolfi;Gleb Shevchuk;Dorsa Sadigh
中科院分区:
其他
文献类型:
--
作者:
Malayandi Palan;Nicholas C. Landolfi;Gleb Shevchuk;Dorsa Sadigh

文献摘要

被引文献

相似文献

我们的目标是准确有效地学习自主机器人的奖励函数。目前解决这个问题的方法包括反向强化学习(IRL),它使用专家演示,以及基于偏好的学习,它迭代地查询用户在轨迹之间的偏好。然而,在机器人技术中,IRL经常会遇到困难,因为它很难获得高质量的演示;相反,基于偏好的学习效率很低,因为它试图从二进制反馈中学习一个连续的高维函数。我们提出了一个新的奖励学习框架,DemPref,它使用演示和偏好查询来学习奖励函数。具体来说,我们(1)使用演示来学习奖励函数空间上的粗糙先验,以减少生成查询的空间的有效大小;以及(2)使用演示来接地(主动)查询生成过程,以提高生成的查询的质量。我们的方法解决了标准的基于偏好的学习方法所面临的效率问题,并且不完全依赖于(可能是低质量的)演示。在数值实验中,我们发现DemPref比标准的基于主动偏好的学习方法更有效。在一项用户研究中,我们将我们的方法与标准的IRL方法进行了比较;我们发现,用户认为使用DemPref训练的机器人在学习他们想要的行为方面更成功,并且更喜欢使用DemPref系统(超过IRL)来训练机器人。
Our goal is to accurately and efficiently learn reward functions for autonomous robots. Current approaches to this problem include inverse reinforcement learning (IRL), which uses expert demonstrations, and preference-based learning, which iteratively queries the user for her preferences between trajectories. In robotics however, IRL often struggles because it is difficult to get high-quality demonstrations; conversely, preference-based learning is very inefficient since it attempts to learn a continuous, high-dimensional function from binary feedback. We propose a new framework for reward learning, DemPref, that uses both demonstrations and preference queries to learn a reward function. Specifically, we (1) use the demonstrations to learn a coarse prior over the space of reward functions, to reduce the effective size of the space from which queries are generated; and (2) use the demonstrations to ground the (active) query generation process, to improve the quality of the generated queries. Our method alleviates the efficiency issues faced by standard preference-based learning methods and does not exclusively depend on (possibly low-quality) demonstrations. In numerical experiments, we find that DemPref is significantly more efficient than a standard active preference-based learning method. In a user study, we compare our method to a standard IRL method; we find that users rated the robot trained with DemPref as being more successful at learning their desired behavior, and preferred to use the DemPref system (over IRL) to train the robot.