Mining Deep Semantic Representations for Scene Classification of High-Resolution Remote Sensing Imagery

Mining Deep Semantic Representations for Scene Classification of High-Resolution Remote Sensing Imagery
复制标题

挖掘高分辨率遥感图像场景分类的深层语义表示

DOI:
10.1109/tbdata.2019.2916880
复制
发表时间:
2020-09-01
影响因子:
7.2
通讯作者:
Zhang, Liangpei
Zhang, Liangpei
中科院分区:
计算机科学2区
文献类型:
--
作者:
Hu, Fan;Xia, Gui-Song;Zhang, Liangpei

文献摘要

被引文献

相似文献

场景分类是高分辨率遥感图像解译中最基本的任务之一。近年来的许多研究表明,能够挖掘图像潜在语义的概率主题模型可以有效地应用于HRRS场景分类中。然而,现有的基于主题模型的方法都是简单地利用底层手工构造的特征来形成语义特征,严重限制了主题模型派生的语义特征的表示能力。为了缓解这一问题,本文提出使用概率潜在语义分析(pLSA)模型构建强大的语义特征,采用预训练的深度卷积神经网络(cnn)作为特征提取器,而不是依赖于手工制作的特征。具体而言,我们开发了两种生成语义特征的方法,即多尺度深度语义表示(MSDS)和多层次深度语义表示(MLDS),通过从不同层提取CNN特征:(1)在MSDS中,pLSA使用从预训练CNN的卷积层提取的多尺度特征来学习最终的语义特征;(2)在MLDS中,我们从预训练的CNN的全连接层中提取不同大小级别的密集采样图像patch的CNN特征,并将pLSA在每一级别学习到的语义特征进行拼接。我们在两个公共HRRS场景数据集上对这两种方法进行了综合评估,并取得了显著的性能改进。突出的结果表明,pLSA模型能够从深度CNN特征中发现具有相当区别性的语义特征。
Scene classification is one of the most fundamental task in interpretation of high-resolution remote sensing (HRRS) images. Many recent works show that the probabilistic topic models which are capable of mining latent semantics of images can be effectively applied to HRRS scene classification. However, the existing approaches based on topic models simply utilize low-level hand-crafted features to form semantic features, which severely limit the representative capability of the semantic features derived from topic models. To alleviate this problem, this paper propose to build powerful semantic features using the probabilistic latent semantic analysis (pLSA) model, by employing the pre-trained deep convolutional neural networks (CNNs) as feature extractors rather than relying on the hand-crafted features. Specifically, we develop two methods to generate semantic features, called multi-scale deep semantic representation (MSDS) and multi-level deep semantic representation (MLDS), by extracting CNN features from different layers: (1) in MSDS, the final semantic features are learned by the pLSA with multi-scale features extracted from the convolutional layer of a pre-trained CNN; (2) in MLDS, we extract CNN features for densely sampled image patches at different size level from the fully-connected layer of a pre-trained CNN, and concatenate the sematic features learned by the pLSA at each level. We comprehensively evaluate the two methods on two public HRRS scene datasets, and achieve significant performance improvement over the state-of-the-art. The outstanding results demonstrate that the pLSA model is capable of discovering considerably discriminative semantic features from the deep CNN features.