A Lightweight Convolutional Neural Network for Hyperspectral Image Classification

A Lightweight Convolutional Neural Network for Hyperspectral Image Classification
复制标题

用于高光谱图像分类的轻量级卷积神经网络

DOI:
10.1109/tgrs.2020.3014313
复制
发表时间:
2021
影响因子:
8.2
通讯作者:
Li Qingquan
Li Qingquan
中科院分区:
工程技术1区
文献类型:
--
作者:
Jia Sen;Lin Zhijie;Xu Meng;Huang Qiang;Zhou Jun;Jia Xiuping;Li Qingquan

文献摘要

被引文献

相似文献

在高光谱图像中,每个像素对应于地球表面的一小块区域,代表了地物的内在特征,可用于土地覆盖的识别。然而,高光谱图像处理面临着一些关键问题,小样本集问题可能是研究中最具挑战性的问题。深度学习也被引入到高光谱图像分类中,该方法已经在许多领域得到了成功的应用。然而,待调整的大量参数与有限的标记样本之间的较大差距会导致场景的过度拟合,从而不可避免地降低DL模型的泛化能力。针对高光谱图像分类中的小样本问题,提出了一种轻量级卷积神经网络(LWCNN)。特别是,首次采用空间-光谱薛定谔特征映射(SSSE)特征提取联合空间-光谱信息,压缩后的维度可以显著减少后续DL模型的参数个数。其次,设计了双尺度卷积(DSC)模块,从一维向量的角度处理SSSE特征(参数个数进一步减少),并利用DSC过程从不同方面获得能够表示数据分布的层次结构描述。随后,通过一种新的双通道融合(BCF)模块对来自DSC各层的特征向量分别进行滤波,该模块能够很好地编码DSC特征中的本征信息和上下文信息。最后,将过滤后的特征连接在一起,并输入到全局平均汇集分类器中,以获得每个类别的预测概率。在三个著名的高光谱图像数据集上的实验结果表明,改进的LWCNN方法在高光谱图像分类任务的效率和稳健性方面都具有优势,并且在标记样本非常有限的情况下优于其他最新的方法(包括基于传统方法和基于DL的方法)。
In the hyperspectral image, each pixel corresponds to a small area on the Earth's surface and represents the intrinsic characteristic of objects, which can be applied for recognition of land covers. Nevertheless, hyperspectral image processing should face some critical issues, and a small sample set problem may be the most challenging one in the research. Deep learning (DL), which has successfully been applied in many fields, has also been introduced for hyperspectral image classification. However, the large gap between the massive parameters to be tuned and limited labeled samples can lead to overfitting scenario, inevitably deteriorating the generalization ability of the DL model. In this article, a lightweight convolutional neural network (LWCNN) is proposed for hyperspectral image classification to mainly tackle the small sample set problem. Especially, spatial-spectral Schroedinger eigenmaps (SSSE) feature extraction is first adopted to obtain the joint spatial-spectral information, and the compressed dimensionality could significantly reduce the number of parameters in the following DL model. Second, a dual-scale convolution (DSC) module is carefully designed to address the SSSE features from a 1-D vector viewpoint (the number of parameters is further decreased), and the DSC procedure is successively employed to obtain the hierarchical structure description that could represent data distribution from different aspects. Subsequently, the feature vectors from all DSC layers are separately filtered by a new bichannel fusion (BCF) module, which could well encode both the intrinsic and contextual information inside DSC features. Finally, the filtered features are concatenated together and imported into a global average pooling classifier to achieve the predicted probability of each category. Experimental results on three famous hyperspectral image data sets illustrate that the developed LWCNN approach is advantageous in both the efficiency and robustness sides for hyperspectral image classification tasks and outperforms other state-of-the-art methods (both traditional-based and DL-based) with very limited labeled samples.