Learning Rotation-Invariant Convolutional Neural Networks for Object Detection in VHR Optical Remote Sensing Images

Learning Rotation-Invariant Convolutional Neural Networks for Object Detection in VHR Optical Remote Sensing Images
复制标题

学习用于 VHR 光学遥感图像中目标检测的旋转不变卷积神经网络

DOI:
10.1109/tgrs.2016.2601622
复制
发表时间:
2016-12-01
影响因子:
8.2
通讯作者:
Han, Junwei
Han, Junwei
中科院分区:
工程技术1区
文献类型:
--
作者:
Cheng, Gong;Zhou, Peicheng;Han, Junwei

文献摘要

被引文献

相似文献

高分辨率光学遥感图像中的目标检测是遥感图像分析面临的一个基本问题。由于强大的特征表示的进步,基于机器学习的目标检测越来越受到关注。虽然存在许多特征表示,但它们中的大多数是手工制作的或基于浅层学习的特征。随着目标检测任务变得越来越具有挑战性,它们的描述能力变得有限甚至贫乏。最近,深度学习算法,特别是卷积神经网络(CNN),在计算机视觉中显示出更强的特征表示能力。尽管自然场景图像取得了进步,但直接使用CNN特征进行光学遥感图像中的物体检测存在问题,因为难以有效处理物体旋转变化的问题。为了解决这个问题,本文提出了一种新颖有效的方法来学习旋转不变CNN(RICNN)模型,以提高目标检测的性能,这是通过在现有CNN架构的基础上引入和学习新的旋转不变层来实现的。然而,与传统CNN模型的训练仅优化多项逻辑回归目标不同,我们的RICNN模型是通过施加正则化约束来优化新的目标函数来训练的,该约束明确地强制训练样本在旋转之前和之后的特征表示相互映射,从而实现旋转不变性。为了便于训练,我们首先训练旋转不变层,然后对整个RICNN网络进行特定领域的微调,以进一步提高性能。在公开的10类目标检测数据集上的综合评价表明了该方法的有效性。
Object detection in very high resolution optical remote sensing images is a fundamental problem faced for remote sensing image analysis. Due to the advances of powerful feature representations, machine-learning-based object detection is receiving increasing attention. Although numerous feature representations exist, most of them are handcrafted or shallow-learning-based features. As the object detection task becomes more challenging, their description capability becomes limited or even impoverished. More recently, deep learning algorithms, especially convolutional neural networks (CNNs), have shown their much stronger feature representation power in computer vision. Despite the progress made in nature scene images, it is problematic to directly use the CNN feature for object detection in optical remote sensing images because it is difficult to effectively deal with the problem of object rotation variations. To address this problem, this paper proposes a novel and effective approach to learn a rotation-invariant CNN (RICNN) model for advancing the performance of object detection, which is achieved by introducing and learning a new rotation-invariant layer on the basis of the existing CNN architectures. However, different from the training of traditional CNN models that only optimizes the multinomial logistic regression objective, our RICNN model is trained by optimizing a new objective function via imposing a regularization constraint, which explicitly enforces the feature representations of the training samples before and after rotating to be mapped close to each other, hence achieving rotation invariance. To facilitate training, we first train the rotation-invariant layer and then domain-specifically fine-tune the whole RICNN network to further boost the performance. Comprehensive evaluations on a publicly available ten-class object detection data set demonstrate the effectiveness of the proposed method.