Medical Image Segmentation via Cascaded Attention Decoding

Medical Image Segmentation via Cascaded Attention Decoding
复制标题

DOI:
10.1109/wacv56688.2023.00616
复制
发表时间:
2023-01
期刊:
2023 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV)
影响因子:
--
通讯作者:
Md Mostafijur Rahman;R. Marculescu
Md Mostafijur Rahman;R. Marculescu
中科院分区:
其他
文献类型:
--
作者:
Md Mostafijur Rahman;R. Marculescu

文献摘要

相似文献

变形金刚在医学图像分割中显示出巨大的前景,因为他们能够通过自我关注来捕捉远程依赖关系。然而,它们缺乏学习像素之间的局部(上下文)关系的能力。以前的工作试图通过在变压器的编码器或解码器模块中嵌入卷积层来克服这一问题,因此有时会得到不一致的特征。为了解决这一问题,我们提出了一种新的基于注意力的解码器,即级联注意力解码器(CASCADE),它利用了分层视觉转换器的多尺度特性。CASCADE包括i)注意门,它融合了特征和跳过连接;ii)卷积注意模块,它通过抑制背景信息来增强远程和局部上下文。我们使用多阶段功能和损耗聚合框架,因为它们的融合速度更快,性能更好。我们的实验表明,级联变压器的性能明显优于最先进的CNN和基于变压器的方法,分别在骰子和MIU分数上获得了5.07%和6.16%的改进。Casade为设计更好的基于注意力的解码器开辟了新的途径。
Transformers have shown great promise in medical image segmentation due to their ability to capture long-range dependencies through self-attention. However, they lack the ability to learn the local (contextual) relations among pixels. Previous works try to overcome this problem by embedding convolutional layers either in the encoder or decoder modules of transformers thus ending up sometimes with inconsistent features. To address this issue, we propose a novel attention-based decoder, namely CASCaded Attention DEcoder (CASCADE), which leverages the multi-scale features of hierarchical vision transformers. CASCADE consists of i) an attention gate which fuses features with skip connections and ii) a convolutional attention module that enhances the long-range and local context by suppressing background information. We use a multi-stage feature and loss aggregation framework due to their faster convergence and better performance. Our experiments demonstrate that transformers with CASCADE significantly outperform state-of-the-art CNN- and transformer-based approaches, obtaining up to 5.07% and 6.16% improvements in DICE and mIoU scores, respectively. CASCADE opens new ways of designing better attention-based decoders.