Quantifying Hypothesis Space Misspecification in Learning From Human–Robot Demonstrations and Physical Corrections

Quantifying Hypothesis Space Misspecification in Learning From Human–Robot Demonstrations and Physical Corrections
复制标题

DOI:
10.1109/tro.2020.2971415
复制
发表时间:
2020-02
影响因子:
7.8
通讯作者:
Andreea Bobu;Andrea V. Bajcsy;J. Fisac;Sampada Deglurkar;A. Dragan
Andreea Bobu;Andrea V. Bajcsy;J. Fisac;Sampada Deglurkar;A. Dragan
中科院分区:
计算机科学1区
文献类型:
--
作者:
Andreea Bobu;Andrea V. Bajcsy;J. Fisac;Sampada Deglurkar;A. Dragan

文献摘要

被引文献

相似文献

人工输入使自主系统能够提高其能力,并实现复杂的行为,否则自动生成具有挑战性。最近的工作集中在机器人如何使用这样的输入,如演示或纠正,以学习预期的目标。这些技术假设人类期望的目标已经存在于机器人的假设空间内。实际上,这种假设往往是不准确的:总是会有这样的情况,即人可能关心机器人不知道的任务方面。没有这些知识,机器人就无法推断出正确的目标。因此,当机器人的假设空间被错误指定时,即使是跟踪目标不确定性的方法也会失败,因为它们会推理哪个假设可能是正确的,而不是任何假设是否正确。在这篇文章中,我们认为机器人应该明确地推理它在给定假设空间的情况下能够解释人类输入的程度,并使用这种情景信心来告知它应该如何融入人类输入。我们展示了我们的方法在7度的自由度机器人操作器在学习两种重要类型的人类输入:演示运动规划任务和物理校正过程中的机器人的任务执行。
The human input has enabled autonomous systems to improve their capabilities and achieve complex behaviors that are otherwise challenging to generate automatically. Recent work focuses on how robots can use such inputs—such as, demonstrations or corrections—to learn intended objectives. These techniques assume that the human‘s desired objective already exists within the robot's hypothesis space. In reality, this assumption is often inaccurate: there will always be situations where the person might care about aspects of the task that the robot does not know about. Without this knowledge, the robot cannot infer the correct objective. Hence, when the robot's hypothesis space is misspecified, even methods that keep track of uncertainty over the objective fail because they reason about which hypothesis might be correct, and not whether any of the hypotheses are correct. In this article, we posit that the robot should reason explicitly about how well it can explain human inputs given its hypothesis space and use that situational confidence to inform how it should incorporate the human input. We demonstrate our method on a 7 degrees-of-freedom robot manipulator in learning from two important types of human inputs: demonstrations of motion planning tasks and physical corrections during the robot's task execution.