HyFormer: Hybrid Grouping-Aggregation Transformer and Wide-Spanning CNN for Hyperspectral Image Super-Resolution

HyFormer: Hybrid Grouping-Aggregation Transformer and Wide-Spanning CNN for Hyperspectral Image Super-Resolution
复制标题

DOI:
10.3390/rs15174131
复制
发表时间:
2023-08
期刊:
Remote. Sens.
影响因子:
--
通讯作者:
Y. Ji;Jingang Shi;Yaping Zhang;Haokun Yang;Yuan Zong;Ling Xu
Y. Ji;Jingang Shi;Yaping Zhang;Haokun Yang;Yuan Zong;Ling Xu
中科院分区:
其他
文献类型:
--
作者:
Y. Ji;Jingang Shi;Yaping Zhang;Haokun Yang;Yuan Zong;Ling Xu

文献摘要

相似文献

高光谱图像的超分辨率是一项具有实际意义和挑战性的任务,因为它需要重建大量的光谱波段。实现出色的重建结果可以极大地有利于后续的下游任务。目前主流的高光谱超分辨率方法主要利用3D卷积神经网络(3D CNN)进行设计。然而,3D CNN中常用的小核大小限制了模型的感受野,使其无法考虑更广泛的上下文信息。虽然可以通过增大核尺寸来扩大感受野,但这会导致模型参数的急剧增加。此外,为自然图像设计的流行视觉变换器不适合处理HSI。这是因为HSI在空间域中表现出稀疏性,这在使用自注意时可能导致显著的计算资源浪费。在本文中,我们设计了一种名为HyFormer的混合架构,它结合了CNN和Transformer的优点,用于高光谱超分辨率。Transformer分支使光谱内交互能够在每个特定波长处捕获细粒度的上下文细节。同时,CNN分支有助于在不同波长之间进行有效的光谱间特征提取,同时保持大的感受野。具体而言,在Transformer分支,我们提出了一种新的分组聚合Transformer(GAT),包括分组自注意力(GSA)和聚合自注意力(阿萨)。GSA用于提取目标的各种细粒度特征,而阿萨用于促进分配给不同通道的异构纹理之间的交互。在CNN分支中,我们提出了一种宽跨度可分离的3D注意力(WSSA)来扩大感受野,同时保持低参数数。在WSSA的基础上,我们构建了一个宽跨度的CNN模块来有效地提取光谱间特征。大量的实验证明了我们的HyFormer的上级性能。
Hyperspectral image (HSI) super-resolution is a practical and challenging task as it requires the reconstruction of a large number of spectral bands. Achieving excellent reconstruction results can greatly benefit subsequent downstream tasks. The current mainstream hyperspectral super-resolution methods mainly utilize 3D convolutional neural networks (3D CNN) for design. However, the commonly used small kernel size in 3D CNN limits the model’s receptive field, preventing it from considering a wider range of contextual information. Though the receptive field could be expanded by enlarging the kernel size, it results in a dramatic increase in model parameters. Furthermore, the popular vision transformers designed for natural images are not suitable for processing HSI. This is because HSI exhibits sparsity in the spatial domain, which can lead to significant computational resource waste when using self-attention. In this paper, we design a hybrid architecture called HyFormer, which combines the strengths of CNN and transformer for hyperspectral super-resolution. The transformer branch enables intra-spectra interaction to capture fine-grained contextual details at each specific wavelength. Meanwhile, the CNN branch facilitates efficient inter-spectra feature extraction among different wavelengths while maintaining a large receptive field. Specifically, in the transformer branch, we propose a novel Grouping-Aggregation transformer (GAT), comprising grouping self-attention (GSA) and aggregation self-attention (ASA). The GSA is employed to extract diverse fine-grained features of targets, while the ASA facilitates interaction among heterogeneous textures allocated to different channels. In the CNN branch, we propose a Wide-Spanning Separable 3D Attention (WSSA) to enlarge the receptive field while keeping a low parameter number. Building upon WSSA, we construct a wide-spanning CNN module to efficiently extract inter-spectra features. Extensive experiments demonstrate the superior performance of our HyFormer.