Mining visual phrases for long-term visual SLAM

Mining visual phrases for long-term visual SLAM
复制标题

DOI:
10.1109/iros.2014.6942552
复制
发表时间:
2014-11
期刊:
2014 IEEE/RSJ International Conference on Intelligent Robots and Systems
影响因子:
--
通讯作者:
Kanji Tanaka;Yuuto Chokushi;Masatoshi Ando
Kanji Tanaka;Yuuto Chokushi;Masatoshi Ando
中科院分区:
其他
文献类型:
--
作者:
Kanji Tanaka;Yuuto Chokushi;Masatoshi Ando

文献摘要

相似文献

我们提出了一种判别性和紧凑的场景描述符,用于单视图位置识别,有助于在熟悉的、半动态的和部分变化的环境中进行长期的视觉SLAM。与流行的词袋场景描述符依赖于矢量量化视觉特征库相比,我们提出的场景描述符基于原始图像数据库(例如可用的视觉体验,其他同事机器人共享的图像,以及网络上公开可用的图像数据),并直接挖掘它来找到视觉短语(VPs),这些视觉短语可以区分和紧凑地解释输入查询/数据库图像。我们的挖掘方法的动机是最近在公共模式发现领域取得的成功——特别是在场景中挖掘公共视觉模式——并且只需要一个可以在不同时间或一天获得的原始图像库。实验结果表明,尽管我们的场景描述符比传统描述符紧凑得多,但它具有相对较高的识别性能。
We propose a discriminative and compact scene descriptor for single-view place recognition that facilitates long-term visual SLAM in familiar, semi-dynamic and partially changing environments. In contrast to popular bag-of-words scene descriptors, which rely on a library of vector quantized visual features, our proposed scene descriptor is based on a library of raw image data (such as an available visual experience, images shared by other colleague robots, and publicly available image data on the web) and directly mine it to find visual phrases (VPs) that discriminatively and compactly explain an input query / database image. Our mining approach is motivated by recent success in the field of common pattern discovery-specifically mining of common visual patterns among scenes-and requires only a single library of raw images that can be acquired at different time or day. Experimental results show that even though our scene descriptor is significantly more compact than conventional descriptors it has a relatively higher recognition performance.