Preference Learning in Assistive Robotics: Observational Repeated Inverse Reinforcement Learning

Preference Learning in Assistive Robotics: Observational Repeated Inverse Reinforcement Learning
复制标题

辅助机器人中的偏好学习:观察重复逆强化学习

DOI:
--
复制
发表时间:
2018
期刊:
Machine Learning in Health Care
影响因子:
--
通讯作者:
L. Riek
L. Riek
中科院分区:
--
文献类型:
--
作者:
Bryce Woodworth;Francesco Ferrari;Teofilo E. Zosa;L. Riek

文献摘要

被引文献

相似文献

随着机器人在日常生活中变得越来越便宜和普遍,对适应用户个性化需求的自适应行为的需求将不断增加。为了实现这一目标,机器人需要通过交互来了解用户的独特偏好。当前的偏好学习技术缺乏在现实的、交互的、不完全信息的环境中推断长期的、任务无关的偏好的能力。为了解决这一差距,我们引入了一种新的偏好推理公式,灵感来自辅助机器人应用程序,其中机器人必须推断这些类型的偏好只基于观察用户在各种任务中的行为。然后,我们提出了一个候选人的推理算法的基础上最大利润的方法,并评估其性能的机器人辅助prehistorian的上下文中。我们发现,该算法学习预测用户的行为方面,因为它是给定的更多的数据,它显示出强大的收敛性能后,少量的迭代。
As robots become more affordable and more common in everyday life, there will be an ever-increasing demand for adaptive behavior that is personalized to the individual needs of users. To accomplish this, robots will need to learn about their users’ unique preferences through interaction. Current preference learning techniques lack the ability to infer long-term, task-independent preferences in realistic, interactive, incomplete-information settings. To address this gap, we introduce a novel preference-inference formulation, inspired by assistive robotics applications, in which a robot must infer these kinds of preferences based only on observing the user’s behavior in various tasks. We then propose a candidate inference algorithm based on maximum-margin methods, and evaluate its performance in the context of robot-assisted prehabilitation. We find that the algorithm learns to predict aspects of the user’s behavior as it is given more data, and that it shows strong convergence properties after a small number of iterations.