Few-Shot Object Detection With Self-Adaptive Global Similarity and Two-Way Foreground Stimulator in Remote Sensing Images

Few-Shot Object Detection With Self-Adaptive Global Similarity and Two-Way Foreground Stimulator in Remote Sensing Images
复制标题

遥感图像中自适应全局相似性和双向前景刺激器的少镜头目标检测

DOI:
10.1109/jstars.2022.3203126
复制
发表时间:
2022
影响因子:
5.5
通讯作者:
Bin Wang
Bin Wang
中科院分区:
工程技术3区
文献类型:
--
作者:
Yuchen Zhang;Bo Zhang;Bin Wang

文献摘要

相似文献

少镜头目标检测的目的是定位和识别潜在的感兴趣的对象,只有通过使用少量的注释数据,它是有益的遥感图像(RSIs)为基础的应用,如城市监测。以往的基于RSIs的少镜头目标检测工作往往试图将支持图像从类无关特征转换为类特定的向量,然后对待检测的查询图像特征进行特征关注操作。然而,这样的方法仍然面临两个关键的挑战:1)它们忽略了支持查询特征的空间相似性,这对于RSI检测是不可或缺的; 2)它们以单向方式执行特征注意操作,这意味着学习的支持查询关系是不对称的。在本文中,为了解决上述挑战,我们设计了一个少镜头对象检测器,它可以快速,准确地推广到看不见的类别,只有少量的数据。拟议的办法包括两个组成部分:1)自适应全局相似度模块,保留内部上下文信息,计算支持图像和查询图像中对象之间的相似度图; 2)双向前景刺激器模块,将相似度图同时应用于支持图像和查询图像的细节嵌入,充分利用支持信息,进一步加强前景对象并削弱无关样本。在DIOR和NWPU VHR-10数据集上进行了实验,实验结果表明,与现有的几种方法相比,该方法具有明显的优越性。
Few-shot object detection aims to localize and recognize potential objects of interest only by using a few annotated data, and it is beneficial for remote sensing images (RSIs) based applications such as urban monitoring. Previous RSIs-based few-shot object detection works often try to convert the support images from class-agnostic features to class-specific vectors, and then perform feature attention operations on query image features to be tested. However, such methods still face two critical challenges: 1) They ignore the spatial similarity of support-query features, which is indispensable for RSIs detection; 2) They perform the feature attention operation in a unidirectional manner, which means that the learned support- query relations are asymmetric. In this paper, to address the challenges above, we design a few-shot object detector, which can quickly and accurately generalize to unseen categories with only a small amount of data. The proposed approach contains two components: 1) the self-adaptive global similarity module that preserves the internal context information to calculate the similarity map between the objects in support and query images, and 2) the two-way foreground stimulator module that can apply the similarity map to the detailed embeddings of support and query images at the same time to make full use of support information, further strengthening the foreground objects and weakening the unconcerned samples. Experiments are conducted on DIOR and NWPU VHR-10 datasets and their results demonstrate the superiority of the proposed method compared with several state-of-the-art methods.