Token Sparsification for Faster Medical Image Segmentation.

Token Sparsification for Faster Medical Image Segmentation.
复制标题

用于更快医学图像分割的令牌稀疏化。

DOI:
10.1007/978-3-031-34048-2_57
复制
发表时间:
2023
期刊:
Information processing in medical imaging : proceedings of the ... conference
影响因子:
--
通讯作者:
Prasanna,Prateek
Prasanna,Prateek
中科院分区:
--
文献类型:
--
作者:
Zhou,Lei;Liu,Huidong;Bae,Joseph;He,Junjun;Samaras,Dimitris;Prasanna,Prateek

文献摘要

相似文献

我们可以使用稀疏令牌进行密集预测吗,例如,分割?虽然令牌稀疏化已被应用于视觉变换器(ViT),以加速分类,它仍然是未知的,如何执行从稀疏令牌分割。为此,我们重新制定分割为asparse编码令牌完成密集解码(SCD)流水线。我们首先从经验上证明,天真地应用现有的分类标记修剪和掩码图像建模(MIM)方法会导致不适当的采样算法和恢复的稠密特征质量低导致训练失败和效率低下。在本文中,我们提出了软顶K令牌修剪(STP)和多层令牌组装(MTA)来解决这些问题。在稀疏编码中,STP使用轻量级子网络预测令牌重要性得分并对topK令牌进行采样。棘手的topK梯度近似通过连续扰动分数分布。在令牌完成中,MTA通过组装稀疏输出令牌和修剪的多层中间令牌来恢复完整的令牌序列。最后的密集解码级与现有的分段解码器兼容,例如,UNETR.实验表明,配备STPandMTA的SCD管道在训练(吞吐量提高120%)和推理(吞吐量提高60.6%)方面都比没有标记修剪的基线快得多,同时保持分割质量。代码可在这里:https://github.com/cvlab-stonybrook/TokenSparse-for-MedSeg.
Can we use sparse tokens for dense prediction, e.g., segmentation?Although token sparsification has been applied to Vision Transformers (ViT) to accelerate classification, it is still unknown how to perform segmentation from sparse tokens. To this end, we reformulate segmentation as asparse encodingtokencompletiondense decoding(SCD) pipeline. We first empirically show that naïvely applying existing approaches from classification token pruning and masked image modeling (MIM) leads to failure and inefficient training caused by inappropriate sampling algorithms and the low quality of the restored dense features. In this paper, we proposeSoft-topK Token Pruning (STP)andMulti-layer Token Assembly (MTA)to address these problems. Insparse encoding,STPpredicts token importance scores with a lightweight sub-network and samples the topK tokens. The intractable topK gradients are approximated through a continuous perturbed score distribution. Intoken completion,MTArestores a full token sequence by assembling both sparse output tokens and pruned multi-layer intermediate ones. The lastdense decodingstage is compatible with existing segmentation decoders, e.g., UNETR. Experiments show SCD pipelines equipped withSTPandMTAare much faster than baselines without token pruning in both training (up to 120% higher throughput) and inference (up to 60.6% higher throughput) while maintaining segmentation quality. Code is available here: https://github.com/cvlab-stonybrook/TokenSparse-for-MedSeg.