Multi-channel and multi-scale mid-level image representation for scene classification

Multi-channel and multi-scale mid-level image representation for scene classification
复制标题

DOI:
10.1117/1.jei.26.2.023018
复制
发表时间:
2017-04
影响因子:
1.1
通讯作者:
Jinfu Yang;Fei Yang;Guanghui Wang;Ming-ai Li
Jinfu Yang;Fei Yang;Guanghui Wang;Ming-ai Li
中科院分区:
计算机科学4区
文献类型:
--
作者:
Jinfu Yang;Fei Yang;Guanghui Wang;Ming-ai Li

文献摘要

被引文献

相似文献

抽象。基于卷积神经网络(CNN)的方法在场景分类中已经获得了最先进的结果。全连通层输出的特征表达了一维的语义信息,但丢失了对象的细节信息和场景类别的空间信息。相反,深度卷积特征已被证明更适合描述对象本身以及图像中对象之间的空间关系。此外,每层特征图在局部邻域内被最大池化,这削弱了全局一致性的不变性,不利于具有高度复杂变化的场景。为了科普上述问题,提出了一种基于预训练CNN特征的无序多通道中级图像表示方法,以提高分类性能。来自FC层和深度卷积层的两个通道的中级图像表示在多尺度级别上被集成。一个总和池的方法也被用来聚合多尺度的中级图像表示,突出的重要性,有利于场景分类的描述符。在SUN397和MIT 67室内数据集上的实验表明,该方法具有良好的分类性能.
Abstract. Convolutional neural network (CNN)-based approaches have received state-of-the-art results in scene classification. Features from the output of fully connected (FC) layers express one-dimensional semantic information but lose the detailed information of objects and the spatial information of scene categories. On the contrary, deep convolutional features have been proved to be more suitable for describing an object itself and the spatial relations among objects in an image. In addition, the feature map from each layer is max-pooled within local neighborhoods, which weakens the invariance of global consistency and is unfavorable to scenes with highly complicated variation. To cope with the above issues, an orderless multi-channel mid-level image representation on pre-trained CNN features is proposed to improve the classification performance. The mid-level image representation of two channels from the FC layer and the deep convolutional layer are integrated at multi-scale levels. A sum pooling approach is also employed to aggregate multi-scale mid-level image representation to highlight the importance of the descriptors beneficial for scene classification. Extensive experiments on SUN397 and MIT 67 indoor datasets demonstrate that the proposed method achieves promising classification performance.