Mining Deep Semantic Representations for Scene Classification of High-Resolution Remote Sensing Imagery
Mining Deep Semantic Representations for Scene Classification of High-Resolution Remote Sensing Imagery
复制标题
挖掘高分辨率遥感图像场景分类的深层语义表示
DOI:
10.1109/tbdata.2019.2916880
复制
发表时间:
2020-09-01
影响因子:
7.2
通讯作者:
Zhang, Liangpei
中科院分区:
文献类型:
--
作者:
Hu, Fan;Xia, Gui-Song;Zhang, Liangpei
Scene classification is one of the most fundamental task in interpretation of high-resolution remote sensing (HRRS) images. Many recent works show that the probabilistic topic models which are capable of mining latent semantics of images can be effectively applied to HRRS scene classification. However, the existing approaches based on topic models simply utilize low-level hand-crafted features to form semantic features, which severely limit the representative capability of the semantic features derived from topic models. To alleviate this problem, this paper propose to build powerful semantic features using the probabilistic latent semantic analysis (pLSA) model, by employing the pre-trained deep convolutional neural networks (CNNs) as feature extractors rather than relying on the hand-crafted features. Specifically, we develop two methods to generate semantic features, called multi-scale deep semantic representation (MSDS) and multi-level deep semantic representation (MLDS), by extracting CNN features from different layers: (1) in MSDS, the final semantic features are learned by the pLSA with multi-scale features extracted from the convolutional layer of a pre-trained CNN; (2) in MLDS, we extract CNN features for densely sampled image patches at different size level from the fully-connected layer of a pre-trained CNN, and concatenate the sematic features learned by the pLSA at each level. We comprehensively evaluate the two methods on two public HRRS scene datasets, and achieve significant performance improvement over the state-of-the-art. The outstanding results demonstrate that the pLSA model is capable of discovering considerably discriminative semantic features from the deep CNN features.