Token Sparsification for Faster Medical Image Segmentation.
Token Sparsification for Faster Medical Image Segmentation.
复制标题
用于更快医学图像分割的令牌稀疏化。
DOI:
10.1007/978-3-031-34048-2_57
复制
发表时间:
2023
期刊:
影响因子:
--
通讯作者:
Prasanna,Prateek
中科院分区:
文献类型:
--
作者:
Zhou,Lei;Liu,Huidong;Bae,Joseph;He,Junjun;Samaras,Dimitris;Prasanna,Prateek
Can we use sparse tokens for dense prediction, e.g., segmentation?Although token sparsification has been applied to Vision Transformers (ViT) to accelerate classification, it is still unknown how to perform segmentation from sparse tokens. To this end, we reformulate segmentation as asparse encodingtokencompletiondense decoding(SCD) pipeline. We first empirically show that naïvely applying existing approaches from classification token pruning and masked image modeling (MIM) leads to failure and inefficient training caused by inappropriate sampling algorithms and the low quality of the restored dense features. In this paper, we proposeSoft-topK Token Pruning (STP)andMulti-layer Token Assembly (MTA)to address these problems. Insparse encoding,STPpredicts token importance scores with a lightweight sub-network and samples the topK tokens. The intractable topK gradients are approximated through a continuous perturbed score distribution. Intoken completion,MTArestores a full token sequence by assembling both sparse output tokens and pruned multi-layer intermediate ones. The lastdense decodingstage is compatible with existing segmentation decoders, e.g., UNETR. Experiments show SCD pipelines equipped withSTPandMTAare much faster than baselines without token pruning in both training (up to 120% higher throughput) and inference (up to 60.6% higher throughput) while maintaining segmentation quality. Code is available here: https://github.com/cvlab-stonybrook/TokenSparse-for-MedSeg.