MoViT: Memorizing Vision Transformers for Medical Image Analysis.

MoViT: Memorizing Vision Transformers for Medical Image Analysis.
复制标题

MoViT:记忆用于医学图像分析的视觉变换器。

DOI:
10.1007/978-3-031-45676-3_21
复制
发表时间:
2024
期刊:
Machine learning in medical imaging. MLMI (Workshop)
影响因子:
--
通讯作者:
Unberath,Mathias
Unberath,Mathias
中科院分区:
--
文献类型:
--
作者:
Shen,Yiqing;Guo,Pengfei;Wu,Jingpu;Huang,Qianqi;Le,Nhat;Zhou,Jinyuan;Jiang,Shanshan;Unberath,Mathias

文献摘要

相似文献

变压器的远程依赖性和卷积神经网络(CNN)图像内容的本地表示的协同作用,由于其互补优势,导致了各种医学图像分析任务的先进架构和性能提高。然而,与CNN相比,变压器需要更多的训练数据,这是由于大量的参数和缺乏电感偏置。对越来越大的数据集的需求仍然存在问题,特别是在医学成像的背景下,其中注释工作和数据保护都导致数据可用性有限。在这项工作中,受人类决策过程的启发,将新的“证据”与先前记忆的“经验”相关联,我们提出了一种简化的视觉transformer Transformer(MoViT),以减轻对大规模数据集的需求,从而成功地训练和部署基于transformer的架构。MoViT利用外部存储器结构在训练阶段缓存历史注意力快照。为了防止过拟合,我们采用了一种创新的内存更新方案,注意时间移动平均,更新存储的外部记忆与历史移动平均。为了加快推理速度,我们设计了一个典型的注意力学习方法,将外部记忆提取成更小的代表子集。我们评估我们的方法在公共组织学图像数据集和内部MRI数据集,证明MoViT应用于各种医学图像分析任务,可以在不同的数据制度优于香草Transformer模型,特别是在只有少量注释数据可用的情况下。更重要的是,MoViT仅用3.0%的训练数据就可以达到ViT的竞争性能。总之,MoViT为Transformer架构提供了一个简单的插件,这可能有助于减少为实现广泛的医学图像分析任务的可接受模型所需的训练数据。
The synergy of long-range dependencies from transformers and local representations of image content from convolutional neural networks (CNNs) has led to advanced architectures and increased performance for various medical image analysis tasks due to their complementary benefits. However, compared with CNNs, transformers require considerably more training data, due to a larger number of parameters and an absence of inductive bias. The need for increasingly large datasets continues to be problematic, particularly in the context of medical imaging, where both annotation efforts and data protection result in limited data availability. In this work, inspired by the human decision-making process of correlating new “evidence” with previously memorized “experience”, we propose a Memorizing Vision Transformer (MoViT) to alleviate the need for large-scale datasets to successfully train and deploy transformer-based architectures. MoViT leverages an external memory structure to cache history attention snapshots during the training stage. To prevent overfitting, we incorporate an innovative memory update scheme, attention temporal moving average, to update the stored external memories with the historical moving average. For inference speedup, we design a prototypical attention learning method to distill the external memory into smaller representative subsets. We evaluate our method on a public histology image dataset and an in-house MRI dataset, demonstrating that MoViT applied to varied medical image analysis tasks, can outperform vanilla transformer models across varied data regimes, especially in cases where only a small amount of annotated data is available. More importantly, MoViT can reach a competitive performance of ViT with only 3.0% of the training data. In conclusion, MoViT provides a simple plug-in for transformer architectures which may contribute to reducing the training data needed to achieve acceptable models for a broad range of medical image analysis tasks.