A Multiview Approach to Learning Articulated Motion Models

A Multiview Approach to Learning Articulated Motion Models
复制标题

DOI:
10.1007/978-3-030-28619-4_30
复制
发表时间:
2017
期刊:
影响因子:
4.4
通讯作者:
Andrea F. Daniele;T. Howard;Matthew R. Walter
Andrea F. Daniele;T. Howard;Matthew R. Walter
中科院分区:
生物学3区
文献类型:
--
作者:
Andrea F. Daniele;T. Howard;Matthew R. Walter

文献摘要

被引文献

相似文献

为了让机器人在家庭和工作场所有效地工作,它们必须能够操纵为人类建造的环境中常见的铰接物体。运动学模型提供了一个简洁的表示这些对象,使故意的,可推广的操纵政策。然而,现有的方法来学习这些模型依赖于视觉观察对象的运动,并受到闭塞和特征稀疏的影响。自然语言描述提供了一种灵活且有效的手段,通过该手段,人类可以以适合于各种不同交互的弱监督方式提供补充信息(例如,演示和远程操作)。在本文中,我们提出了一个多模态学习框架,它结合了视觉和语言的信息,在现场获得的结构和参数,定义的运动模型的关节连接的对象估计。视觉信号采用RGB-D图像流的形式,该图像流在无准备的场景中机会性地捕获对象运动。伴随的运动的自然语言描述构成语言信号。我们使用概率图形模型,自然语言的描述,其所指的运动学运动的基础语言信息建模。通过利用视觉和语言观察的互补性,我们的方法为各种多部分对象推断出正确的运动学模型,而这些对象是以前最先进的仅视觉系统失败的。我们在由各种家庭对象组成的数据集上评估了我们的多模态学习框架,并证明了模型准确性在仅视觉基线上的改进。
In order for robots to operate effectively in homes and workplaces, they must be able to manipulate the articulated objects common within environments built for and by humans. Kinematic models provide a concise representation of these objects that enable deliberate, generalizable manipulation policies. However, existing approaches to learning these models rely upon visual observations of an object’s motion, and are subject to the effects of occlusions and feature sparsity. Natural language descriptions provide a flexible and efficient means by which humans can provide complementary information in a weakly supervised manner suitable for a variety of different interactions (e.g., demonstrations and remote manipulation). In this paper, we present a multimodal learning framework that incorporates both vision and language information acquired in situ to estimate the structure and parameters that define kinematic models of articulated objects. The visual signal takes the form of an RGB-D image stream that opportunistically captures object motion in an unprepared scene. Accompanying natural language descriptions of the motion constitute the linguistic signal. We model linguistic information using a probabilistic graphical model that grounds natural language descriptions to their referent kinematic motion. By exploiting the complementary nature of the vision and language observations, our method infers correct kinematic models for various multiple-part objects on which the previous state-of-the-art, visual-only system fails. We evaluate our multimodal learning framework on a dataset comprised of a variety of household objects, and demonstrate aimprovement in model accuracy over the vision-only baseline.