Multi-scale Hierarchical Vision Transformer with Cascaded Attention Decoding for Medical Image Segmentation

Multi-scale Hierarchical Vision Transformer with Cascaded Attention Decoding for Medical Image Segmentation
复制标题

DOI:
10.48550/arxiv.2303.16892
复制
发表时间:
2023-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Md Mostafijur Rahman;R. Marculescu
Md Mostafijur Rahman;R. Marculescu
中科院分区:
其他
文献类型:
--
作者:
Md Mostafijur Rahman;R. Marculescu

文献摘要

相似文献

变形金刚在医学图像分割中取得了巨大的成功。然而,由于潜在的单尺度自我注意(SA)机制,变压器可能表现出有限的泛化能力。在本文中,我们通过引入多尺度分层视觉转换器(MERIT)骨干网络来解决这一问题,该网络通过在多个尺度上计算SA来提高模型的泛化能力。我们还加入了一个基于注意力的解码器,即级联注意力解码(CASCADE),以进一步细化由优点产生的多阶段特征。最后,我们介绍了一种有效的多阶段特征混合损失聚集(变异)方法,以通过隐式集成来更好地训练模型。我们在两个广泛使用的医学图像分割基准(即Synapse多器官,ACDC)上的实验表明,该方法比最先进的方法具有更好的性能。我们的优点架构和突变丢失聚合可用于下游医学图像和语义分割任务。
Transformers have shown great success in medical image segmentation. However, transformers may exhibit a limited generalization ability due to the underlying single-scale self-attention (SA) mechanism. In this paper, we address this issue by introducing a Multi-scale hiERarchical vIsion Transformer (MERIT) backbone network, which improves the generalizability of the model by computing SA at multiple scales. We also incorporate an attention-based decoder, namely Cascaded Attention Decoding (CASCADE), for further refinement of multi-stage features generated by MERIT. Finally, we introduce an effective multi-stage feature mixing loss aggregation (MUTATION) method for better model training via implicit ensembling. Our experiments on two widely used medical image segmentation benchmarks (i.e., Synapse Multi-organ, ACDC) demonstrate the superior performance of MERIT over state-of-the-art methods. Our MERIT architecture and MUTATION loss aggregation can be used with downstream medical image and semantic segmentation tasks.