Learning Rotation-Invariant and Fisher Discriminative Convolutional Neural Networks for Object Detection

Learning Rotation-Invariant and Fisher Discriminative Convolutional Neural Networks for Object Detection
复制标题

学习用于目标检测的旋转不变和 Fisher 判别卷积神经网络

DOI:
10.1109/tip.2018.2867198
复制
发表时间:
2019-01-01
影响因子:
10.6
通讯作者:
Xu, Dong
Xu, Dong
中科院分区:
计算机科学1区
文献类型:
--
作者:
Cheng, Gong;Han, Junwei;Xu, Dong

文献摘要

被引文献

相似文献

最近,由于卷积神经网络(CNN)学习到的强大特征,目标检测的性能得到了显著的提高。尽管已经取得了显著的成功,但目标检测仍然面临着几个主要挑战,包括目标轮换、类内多样性和类间相似性,这通常会降低目标检测的性能。为了解决这些问题,我们建立了现有的目标检测系统,并提出了一种简单而有效的方法来训练旋转不变和Fisher判别的CNN模型,以进一步提高目标检测的性能。这是通过优化一个新的目标函数来实现的,该目标函数显式地将旋转不变正则化和Fisher判别正则化强加于CNN特征。具体地说,第一正则化器强制旋转前后的训练样本的CNN特征表示彼此紧密映射,以实现旋转不变性。第二种正则化算法约束CNN特征具有较小的类内离散度和较大的类间分离度。我们在四种流行的目标检测框架下实现了我们的方法,包括Region-CNN(R-CNN)、Fast R-CNN、Fast R-CNN和R-FCN。在实验中,我们在Pascal VOC 2007和2012数据集以及公开可用的航空图像数据集上对所提出的方法进行了全面的评估。我们提出的方法比现有的基线方法性能更好,并达到了最先进的结果。
The performance of object detection has recently been significantly improved due to the powerful features learnt through convolutional neural networks (CNNs). Despite the remarkable success, there are still several major challenges in object detection, including object rotation, within-class diversity, and between-class similarity, which generally degenerate object detection performance. To address these issues, we build up the existing state-of-the-art object detection systems and propose a simple but effective method to train rotation-invariant and Fisher discriminative CNN models to further boost object detection performance. This is achieved by optimizing a new objective function that explicitly imposes a rotation-invariant regularizer and a Fisher discrimination regularizer on the CNN features. Specifically, the first regularizer enforces the CNN feature representations of the training samples before and after rotation to be mapped closely to each other in order to achieve rotation-invariance. The second regularizer constrains the CNN features to have small within-class scatter but large between-class separation. We implement our proposed method under four popular object detection frameworks, including region-CNN (R-CNN), Fast R- CNN, Faster R- CNN, and R- FCN. In the experiments, we comprehensively evaluate the proposed method on the PASCAL VOC 2007 and 2012 data sets and a publicly available aerial image data set. Our proposed methods outperform the existing baseline methods and achieve the state-of-the-art results.