Revisiting Monocular Satellite Pose Estimation With Transformer

Revisiting Monocular Satellite Pose Estimation With Transformer
复制标题

DOI:
10.1109/taes.2022.3161605
复制
发表时间:
2022-10
影响因子:
4.4
通讯作者:
Zi Wang;Zhuo Zhang;Xiaoliang Sun;Zhang Li;Qifeng Yu
Zi Wang;Zhuo Zhang;Xiaoliang Sun;Zhang Li;Qifeng Yu
中科院分区:
计算机科学2区
文献类型:
--
作者:
Zi Wang;Zhuo Zhang;Xiaoliang Sun;Zhang Li;Qifeng Yu

文献摘要

相似文献

卷积神经网络(Convolutional neural networks, cnn)被应用于单眼卫星位姿估计中,取得了优于传统方法的性能。然而,现有的基于cnn的方法存在着对纹理的偏向、对绝对距离的间接描述以及缺乏远程依赖建模等问题。这些因素限制了基于cnn的方法的通用性。在变压器模型取得显著成果的启发下,本文采用变压器块对单幅RGB图像进行卫星位姿估计,提出了一种高效的单眼卫星位姿估计方法。首先,我们基于一组关键点设计了一个有效的卫星表示模型。然后,考虑单眼卫星姿态估计的特点,构建端到端关键点集预测网络,构建二部损失函数;进一步,我们改进了主干结构,以获得高质量的特征提取。在公共基准数据集上的实验结果表明,该方法仅使用合成训练数据,在合成测试集和真实测试集上分别获得第二和第三名。我们还证明,在我们的比较中,我们的关键点预测器花费的时间是第一种方法的一半,因此在速度和准确性之间实现了比现有方法更好的权衡。
Convolutional neural networks (CNNs) have been adopted in monocular satellite pose estimation and achieve superior performance over traditional methods. However, existing CNN-based methods suffer from bias toward texture, indirect description of absolute distance, and lack of long-range dependence modeling. Such factors limit the generalizability of CNN-based methods. Motivated by the striking achievements of transformer models, this article adopts transformer blocks for satellite pose estimation from a single RGB image, proposing an efficient monocular satellite pose estimation method. First, we design an effective satellite representation model based on a set of keypoints. Then, considering monocular satellite pose estimation characteristics, we construct an end-to-end keypoint-set prediction network and build the bipartite loss function. Further, we improve the backbone structure for high-quality feature extraction. Experimental results on a public benchmark dataset indicate that the proposed method achieves second and third place on the synthetic and real test sets, respectively, using only synthetic training data. We also demonstrate that our keypoint predictor takes half as much time as the first-placed method in our comparison, and therefore achieves a better tradeoff between speed and accuracy than existing approaches.