SwinNet: Swin Transformer Drives Edge-Aware RGB-D and RGB-T Salient Object Detection

SwinNet: Swin Transformer Drives Edge-Aware RGB-D and RGB-T Salient Object Detection
复制标题

DOI:
10.1109/tcsvt.2021.3127149
复制
发表时间:
2022-07-01
影响因子:
8.4
通讯作者:
Xiao, Yun
Xiao, Yun
中科院分区:
工程技术1区
文献类型:
--
作者:
Liu, Zhengyi;Tan, Yacheng;Xiao, Yun

文献摘要

被引文献

相似文献

卷积神经网络(CNN)擅长提取某些感受野内的纹理特征,而transformers可以对全局长程依赖特征进行建模。Swin Transformer吸收了Transformer的优点和CNN的优点,表现出较强的特征表示能力。在此基础上,我们提出了一种粗糙模态融合模型SwinNet,用于RGB-D和RGB-T显著对象检测。该方法以Swin Transformer为驱动,提取层次特征,以注意力机制为辅助,弥合两种模态之间的差距,以边缘信息为引导,锐化显著对象的轮廓。具体而言,两码流Swin Transformer编码器首先提取多模态特征,然后提出空间对齐和通道重校准模块,优化层内跨模态特征。边缘引导的解码器在边缘特征的指导下实现跨模态融合,以清晰模糊边界。该模型在RGB-D和RGB-T数据集上的表现优于最先进的模型,表明它提供了对跨模态互补任务的更多洞察。
Convolutional neural networks (CNNs) are good at extracting contexture features within certain receptive fields, while transformers can model the global long-range dependency features. By absorbing the advantage of transformer and the merit of CNN, Swin Transformer shows strong feature representation ability. Based on it, we propose a crass-modality fusion model, SwinNet, for RGB-D and RGB-T salient object detection. It is driven by Swin Transformer to extract the hierarchical features, boosted by attention mechanism to bridge the gap between two modalities, and guided by edge information to sharp the contour of salient object. To be specific, two-stream Swin Transformer encoder first extracts multi-modality features, and then spatial alignment and channel re-calibration module is presented to optimize intra-level cross-modality features. To clarify the fumy boundary, edge-guided decoder achieves interlevel cross-modality fusion under the guidance of edge features. The proposed model outperforms the state-of-the-art models on RGB-D and RGB-T datasets, showing that it provides more insight into the cross-modality complementarily task.