Towards Real-Time Segmentation on the Edge

Towards Real-Time Segmentation on the Edge
复制标题

DOI:
10.1609/aaai.v37i2.25232
复制
发表时间:
2023-06
期刊:
--
影响因子:
--
通讯作者:
Yanyu Li;Changdi Yang;Pu Zhao;Geng Yuan;Wei Niu;Jiexiong Guan;Hao Tang;Minghai Qin;Qing Jin;Bin Ren;Xue Lin;Yanzhi Wang
Yanyu Li;Changdi Yang;Pu Zhao;Geng Yuan;Wei Niu;Jiexiong Guan;Hao Tang;Minghai Qin;Qing Jin;Bin Ren;Xue Lin;Yanzhi Wang
中科院分区:
其他
文献类型:
--
作者:
Yanyu Li;Changdi Yang;Pu Zhao;Geng Yuan;Wei Niu;Jiexiong Guan;Hao Tang;Minghai Qin;Qing Jin;Bin Ren;Xue Lin;Yanzhi Wang

文献摘要

相似文献

实时分割的研究主要集中在桌面GPU上。然而,自动驾驶和许多其他应用都依赖于边缘的实时分割,而目前的技术还远远没有实现这一目标。此外,视觉转换器的最新进展也启发我们重新设计密集预测任务的网络架构。在这项工作中,我们提出了联合收割机结合自注意块与轻量级卷积,以形成新的构建块,并采用延迟约束,以搜索一个有效的子网络。我们根据生成的架构配置及其在移动的设备上测量的延迟来训练MLP延迟模型,以便我们可以预测搜索阶段的延迟。据我们所知,我们是第一个在Cityscapes上实现超过74%的mIoU,并在移动的GPU上通过现成的手机进行半实时推理(超过15 FPS)。
The research in real-time segmentation mainly focuses on desktop GPUs. However, autonomous driving and many other applications rely on real-time segmentation on the edge, and current arts are far from the goal. In addition, recent advances in vision transformers also inspire us to re-design the network architecture for dense prediction task. In this work, we propose to combine the self attention block with lightweight convolutions to form new building blocks, and employ latency constraints to search an efficient sub-network. We train an MLP latency model based on generated architecture configurations and their latency measured on mobile devices, so that we can predict the latency of subnets during search phase. To the best of our knowledge, we are the first to achieve over 74% mIoU on Cityscapes with semi-real-time inference (over 15 FPS) on mobile GPU from an off-the-shelf phone.