Pruning Parameterization with Bi-level Optimization for Efficient Semantic Segmentation on the Edge

Pruning Parameterization with Bi-level Optimization for Efficient Semantic Segmentation on the Edge
复制标题

DOI:
10.1109/cvpr52729.2023.01478
复制
发表时间:
2023-06
期刊:
2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)
影响因子:
--
通讯作者:
Changdi Yang;Pu Zhao;Yanyu Li;Wei Niu;Jiexiong Guan;Hao Tang;Minghai Qin;Bin Ren;Xue Lin;Yanzhi Wang
Changdi Yang;Pu Zhao;Yanyu Li;Wei Niu;Jiexiong Guan;Hao Tang;Minghai Qin;Bin Ren;Xue Lin;Yanzhi Wang
中科院分区:
其他
文献类型:
--
作者:
Changdi Yang;Pu Zhao;Yanyu Li;Wei Niu;Jiexiong Guan;Hao Tang;Minghai Qin;Bin Ren;Xue Lin;Yanzhi Wang

文献摘要

相似文献

随着边缘设备的日益普及,自动驾驶和许多其他应用需要在边缘实现实时分割。视觉转换器(ViTs)在许多视觉任务中表现出相当强的效果。然而,具有全注意力机制的ViT通常消耗大量的计算资源,导致边缘设备上的真实的实时推理的困难。在本文中,我们的目标是以更少的计算和更快的推理速度获得ViTs,以促进边缘设备上语义分割的密集预测。为了实现这一点,我们提出了一个修剪参数化方法来制定语义分割的修剪问题。然后,我们采用了一个双层优化方法来解决这个问题的隐式梯度的帮助。我们的实验结果表明,我们可以实现38.9 mIoU的ADE 20 K瓦尔与速度为56.5 FPS的三星S21,这是最高的mIoU在相同的计算约束下,实时推理。
With the ever-increasing popularity of edge devices, it is necessary to implement real-time segmentation on the edge for autonomous driving and many other applications. Vision Transformers (ViTs) have shown considerably stronger results for many vision tasks. However, ViTs with the fullattention mechanism usually consume a large number of computational resources, leading to difficulties for real- time inference on edge devices. In this paper, we aim to derive ViTs with fewer computations and fast inference speed to facilitate the dense prediction of semantic segmentation on edge devices. To achieve this, we propose a pruning parameterization method to formulate the pruning problem of semantic segmentation. Then we adopt a bi-level optimization method to solve this problem with the help of implicit gradients. Our experimental results demonstrate that we can achieve 38.9 mIoU on ADE20K val with a speed of 56.5 FPS on Samsung S21, which is the highest mIoU under the same computation constraint with real-time inference.