GTCaR: Graph Transformer for Camera Re-localization

GTCaR: Graph Transformer for Camera Re-localization
复制标题

DOI:
10.1007/978-3-031-20080-9_14
复制
发表时间:
2022
期刊:
--
影响因子:
--
通讯作者:
Xinyi Li;Haibin Ling
Xinyi Li;Haibin Ling
中科院分区:
其他
文献类型:
--
作者:
Xinyi Li;Haibin Ling

文献摘要

相似文献

摄像机重新定位或绝对姿态回归是许多计算机视觉任务的核心,例如视觉里程计、运动结构(SFM)和SLAM。本文提出了一种以图形转换器为主干的神经网络方法,称为GTCaR(GraphTransformer for CameraRe-Location),用于解决多视点摄像机的再定位问题。与以往的姿态回归主要由光度一致性指导的工作不同,GTCAR有效地将图像特征、摄像机姿态信息和帧间相对摄像机运动融合到编码的图属性中。此外,GTCaR被训练成图形一致性和姿态精度相结合的方式,从而产生了显著更高的计算效率。通过利用具有边缘特征的图变换器层和启用邻接张量,GTCaR动态地捕获全局注意力,从而赋予姿势图以进化的结构,以实现更好的稳健性和准确性。此外,可选的时间转换器层主动增强了顺序输入的时空帧间关系。对建议的网络在各种公共基准上的评估表明,GTCaR的表现优于最先进的方法。
Camera re-localization or absolute pose regression is the centerpiece in numerous computer vision tasks such as visual odometry, structure from motion (SfM) and SLAM. In this paper we propose a neural network approach with a graph Transformer backbone, namelyGTCaR(GraphTransformer forCameraRe-localization), to address the multi-view camera re-localization problem. In contrast with prior work where the pose regression is mainly guided by photometric consistency, GTCaR effectively fuses the image features, camera pose information and inter-frame relative camera motions into encoded graph attributes. Moreover, GTCaR is trained towards the graph consistency and pose accuracy combined instead, yielding significantly higher computational efficiency. By leveraging graph Transformer layers with edge features and enabling the adjacency tensor, GTCaR dynamically captures the global attention and thus endows the pose graph with evolving structures to achieve improved robustness and accuracy. In addition, optional temporal Transformer layers actively enhance the spatiotemporal inter-frame relation for sequential inputs. Evaluation of the proposed network on various public benchmarks demonstrates that GTCaR outperforms state-of-the-art approaches.