Mining DCNN landmarks for long-term visual SLAM

Mining DCNN landmarks for long-term visual SLAM
复制标题

DOI:
10.1109/robio.2016.7866383
复制
发表时间:
2016-12
期刊:
2016 IEEE International Conference on Robotics and Biomimetics (ROBIO)
影响因子:
--
通讯作者:
Taisho Tsukamoto;Kanji Tanaka
Taisho Tsukamoto;Kanji Tanaka
中科院分区:
其他
文献类型:
--
作者:
Taisho Tsukamoto;Kanji Tanaka

文献摘要

相似文献

在熟悉的、半动态的和部分变化的环境中的长期视觉SLAM是机器人研究的一个重要领域。我们面临的主要问题是如何有区别地和紧凑地描述一个场景,这两个都是必要的,以科普外观和大量的视觉信息的变化。在本研究中,我们通过挖掘视觉经验来解决上述问题。我们的策略是挖掘一个原始视觉图像库,称为视觉体验,以找到相关的视觉模式,有效地解释输入场景。从实践的角度来看,我们的工作提供了三个主要的贡献比以前的工作。首先,它是深度卷积神经网络(DCNN)的区分性视觉特征首次应用于视觉地标挖掘任务。其次,我们展示了如何将高维DCNN特征解释为视觉词的紧凑语义表示。第三,我们证明了我们的方法可以将具有任何特征(包括DCNN特征)的场景描述任务转化为挖掘视觉体验的任务。在一个具有挑战性的跨域视觉场所识别实验上验证了该方法的有效性。
Long-term visual SLAM, in familiar, semi-dynamic, and partially changing environments is an important area of research in robotics. The main problem we faced is the question of how to describe a scene discriminatively and compactly-both of which are necessary in order to cope with changes in appearance and a large amount of visual information. In this study, we address the above issues by mining visual experience. Our strategy is to mine a library of raw visual images, termed visual experience, to find the relevant visual patterns to effectively explain the input scene. From a practical point of view, our work offers three main contributions over the previous work. First, it is the first application of discriminative visual features from deep convolutional neural networks (DCNN) to the task of visual landmark mining. Second, we show how to interpret a high-dimensional DCNN feature to a compact semantic representation of visual word. Third, we show that our approach can turn the scene description task with any feature (including the DCNN feature) into the task of mining visual experience. Experiments on a challenging cross-domain visual place recognition validate efficacy of the proposed approach.