Multimodal estimation and communication of latent semantic knowledge for robust execution of robot instructions

Multimodal estimation and communication of latent semantic knowledge for robust execution of robot instructions
复制标题

DOI:
10.1177/0278364920917755
复制
发表时间:
2020-06-05
影响因子:
9.2
通讯作者:
Paul, Rohan
Paul, Rohan
中科院分区:
计算机科学2区
文献类型:
--
作者:
Arkin, Jacob;Park, Daehyung;Paul, Rohan

文献摘要

被引文献

相似文献

本文的目标是使机器人能够在部分可观察的环境中按照人类指令执行健壮的任务。机器人解释和执行命令的能力从根本上取决于其语义世界知识。通常,机器人使用摄像机或激光雷达等外感受传感器来检测工作空间中的实体并推断它们的视觉属性和空间关系。然而,语义世界的属性通常是视觉上无法感知的。我们使用非外感受的形式,包括物理本体感受,事实描述和领域知识的机制,推断对象的语义属性。我们引入了一个概率模型,该模型将语言知识与视觉和触觉观察融合到对潜在世界属性的累积信念中,以推断指令的含义并以对错误、噪音或矛盾证据鲁棒的方式执行指示的任务。此外,我们提供了一种方法,允许机器人沟通的知识不和谐回人类作为一种手段,纠正操作员的世界模型中的错误。最后,我们提出了一个有效的框架,预计可能的语言互动和推断相关的接地为当前的世界状态,从而引导语言理解和生成。我们目前的实验操纵器的任务,需要推断部分观察到的语义属性,并评估我们的框架的能力,利用表达的信息和知识库,以促进收敛,并生成声明的声明声明的事实,观察到不一致的机器人的估计对象属性。
The goal of this article is to enable robots to perform robust task execution following human instructions in partially observable environments. A robot's ability to interpret and execute commands is fundamentally tied to its semantic world knowledge. Commonly, robots use exteroceptive sensors, such as cameras or LiDAR, to detect entities in the workspace and infer their visual properties and spatial relationships. However, semantic world properties are often visually imperceptible. We posit the use of non-exteroceptive modalities including physical proprioception, factual descriptions, and domain knowledge as mechanisms for inferring semantic properties of objects. We introduce a probabilistic model that fuses linguistic knowledge with visual and haptic observations into a cumulative belief over latent world attributes to infer the meaning of instructions and execute the instructed tasks in a manner robust to erroneous, noisy, or contradictory evidence. In addition, we provide a method that allows the robot to communicate knowledge dissonance back to the human as a means of correcting errors in the operator's world model. Finally, we propose an efficient framework that anticipates possible linguistic interactions and infers the associated groundings for the current world state, thereby bootstrapping both language understanding and generation. We present experiments on manipulators for tasks that require inference over partially observed semantic properties, and evaluate our framework's ability to exploit expressed information and knowledge bases to facilitate convergence, and generate statements to correct declared facts that were observed to be inconsistent with the robot's estimate of object properties.