Towards Transferring Human Preferences from Canonical to Actual Assembly Tasks

Towards Transferring Human Preferences from Canonical to Actual Assembly Tasks
复制标题

DOI:
10.1109/ro-man53752.2022.9900872
复制
发表时间:
2021-11
期刊:
2022 31st IEEE International Conference on Robot and Human Interactive Communication (RO-MAN)
影响因子:
--
通讯作者:
Heramb Nemlekar;Runyu Guan;Guanyang Luo;Satyandra K. Gupta;S. Nikolaidis
Heramb Nemlekar;Runyu Guan;Guanyang Luo;Satyandra K. Gupta;S. Nikolaidis
中科院分区:
其他
文献类型:
--
作者:
Heramb Nemlekar;Runyu Guan;Guanyang Luo;Satyandra K. Gupta;S. Nikolaidis

文献摘要

被引文献

相似文献

为了帮助人类用户根据他们的个人偏好在装配任务,机器人通常需要在给定的任务中的用户演示。然而,在实际装配任务中提供演示可能是乏味和耗时的。我们的论点是,我们可以学习用户在实际装配任务的偏好,从他们的示范,在一个典型的任务。受人类运动经济学先前工作的启发,我们建议将用户偏好表示为抽象任务不可知特征(如用户所需的运动和身心努力)的线性奖励函数。对于每个用户,我们从他们在规范任务中的演示中学习奖励函数的权重,并使用学习到的权重来预测他们在实际组装任务中的动作;在实际任务中没有任何用户演示。我们评估我们提出的方法在模型飞机装配研究,并表明,偏好可以有效地从规范转移到实际的装配任务,使机器人能够预测用户的行动。
To assist human users according to their individual preference in assembly tasks, robots typically require user demonstrations in the given task. However, providing demonstrations in actual assembly tasks can be tedious and time-consuming. Our thesis is that we can learn the preference of users in actual assembly tasks from their demonstrations in a representative canonical task. Inspired by prior work in economy of human movement, we propose to represent user preferences as a linear reward function over abstract task-agnostic features, such as movement and physical and mental effort required by the user. For each user, we learn the weights of the reward function from their demonstrations in a canonical task and use the learned weights to anticipate their actions in the actual assembly task; without any user demonstrations in the actual task. We evaluate our proposed method in a model-airplane assembly study and show that preferences can be effectively transferred from canonical to actual assembly tasks, enabling robots to anticipate user actions.