A visual language model for estimating object pose and structure in a generative visual domain

A visual language model for estimating object pose and structure in a generative visual domain
复制标题

用于估计生成视觉领域中的物体姿态和结构的视觉语言模型

DOI:
10.1109/icra.2011.5980161
复制
发表时间:
2011
期刊:
2011 IEEE International Conference on Robotics and Automation
影响因子:
--
通讯作者:
J. Siskind
J. Siskind
中科院分区:
--
文献类型:
--
作者:
Siddharth Narayanaswamy;Andrei Barbu;J. Siskind

文献摘要

被引文献

相似文献

我们通过类比与人类语言的生成性质相比,我们呈现了视觉对象的生成域。正如音素和单词的小库存以语法方式结合在一起,以产生无数有效的单词和话语一样,一个小的物理零件以语法方式结合在一起,以产生无数有效的组件。我们将语言模型的概念从语音识别到此视觉领域的概念,以类似地提高识别过程的性能,而不是仅将识别器应用于组件而可能发生的事情。与人类语言的无上下文模型不同,我们的视觉语言模型对上下文敏感,并将其作为随机约束 - 满足问题进行表述。与人类语言的情况不同,所有组成部分都是可观察到的,我们的方法涉及遮挡,尽管不可观察到的组成部分,但仍成功恢复对象结构。我们使用集成的机器人系统演示了我们的系统,用于拆卸结构,该结构在存在噪声特征探测器的情况下执行与语言模型一致的全景重建。
We present a generative domain of visual objects by analogy to the generative nature of human language. Just as small inventories of phonemes and words combine in a grammatical fashion to yield myriad valid words and utterances, a small inventory of physical parts combine in a grammatical fashion to yield myriad valid assemblies. We apply the notion of a language model from speech recognition to this visual domain to similarly improve the performance of the recognition process over what would be possible by only applying recognizers to the components. Unlike the context-free models for human language, our visual language models are context sensitive and formulated as stochastic constraint-satisfaction problems. And unlike the situation for human language where all components are observable, our methods deal with occlusion, successfully recovering object structure despite unobservable components. We demonstrate our system with an integrated robotic system for disassembling structures that performs whole-scene reconstruction consistent with a language model in the presence of noisy feature detectors.