Transformer guided geometry model for flow-based unsupervised visual odometry

Transformer guided geometry model for flow-based unsupervised visual odometry
复制标题

用于基于流的无监督视觉里程计的变压器引导几何模型

DOI:
10.1007/s00521-020-05545-8
复制
发表时间:
2021-01-02
影响因子:
6
通讯作者:
Li, Wanqing
Li, Wanqing
中科院分区:
计算机科学3区
文献类型:
--
作者:
Li, Xiangyu;Hou, Yonghong;Li, Wanqing

文献摘要

被引文献

相似文献

现有的无监督视觉里程计(VO)方法要么匹配成对的图像或整合的时间信息,使用递归神经网络在一个长序列的图像。它们要么不准确,训练耗时或误差累积。在本文中,我们提出了一种方法,包括两个相机姿态估计处理成对图像和一个短序列的图像,分别从信息。对于图像序列,采用类似于变换器的结构在局部时间窗口上建立几何模型,称为基于变换器的辅助姿态估计器(TAPE)。同时,提出了一种流到流姿态估计器(F2FPE),以利用成对图像之间的关系。这两个估计量通过简单而有效的训练一致性损失来约束。实证评估表明,该方法优于最先进的无监督学习的方法,由一个很大的利润,并执行监督和传统的KITTI和马拉加数据集。
Existing unsupervised visual odometry (VO) methods either match pairwise images or integrate the temporal information using recurrent neural networks over a long sequence of images. They are either not accurate, time-consuming in training or error accumulative. In this paper, we propose a method consisting of two camera pose estimators that deal with the information from pairwise images and a short sequence of images, respectively. For image sequences, a transformer-like structure is adopted to build a geometry model over a local temporal window, referred to as transformer-based auxiliary pose estimator (TAPE). Meanwhile, a flow-to-flow pose estimator (F2FPE) is proposed to exploit the relationship between pairwise images. The two estimators are constrained through a simple yet effective consistency loss in training. Empirical evaluation has shown that the proposed method outperforms the state-of-the-art unsupervised learning-based methods by a large margin and performs comparably to supervised and traditional ones on the KITTI and Malaga dataset.