FASSD-Net: Fast and Accurate Real-Time Semantic Segmentation for Embedded Systems

FASSD-Net: Fast and Accurate Real-Time Semantic Segmentation for Embedded Systems
复制标题

FASSD-Net:快速准确的嵌入式系统实时语义分割

DOI:
10.1109/tits.2021.3127553
复制
发表时间:
2022-09
影响因子:
8.5
通讯作者:
L. Rosas-Arias;Gibran Benitez-Garcia;J. Portillo-Portillo-J.-Portillo-Portillo-1400531221;J. Olivares-Mercado;G. Sánchez-Pérez;Keiji Yanai
L. Rosas-Arias;Gibran Benitez-Garcia;J. Portillo-Portillo-J.-Portillo-Portillo-1400531221;J. Olivares-Mercado;G. Sánchez-Pérez;Keiji Yanai
中科院分区:
工程技术1区
文献类型:
--
作者:
L. Rosas-Arias;Gibran Benitez-Garcia;J. Portillo-Portillo-J.-Portillo-Portillo-1400531221;J. Olivares-Mercado;G. Sánchez-Pérez;Keiji Yanai

文献摘要

相似文献

最近的实时语义分割工作,从密集的深度神经网络中删除或使用轻型解码器,以实现快速推理速度。这种策略有助于实现实时性能;然而,与非实时方法相比,准确性大大降低。在本文中,我们介绍了两个关键模块,旨在设计一个高性能的解码器,用于实时语义分割,这也减少了实时和非实时网络之间的准确性差距。第一个模块,扩张非对称金字塔融合(DAPF),旨在增加最后一级编码器顶部的感受野,获得更丰富的上下文特征。第二个模块,多分辨率扩展非对称(MDA)模块,融合和细化来自网络早期和更深阶段的多尺度特征图的细节和上下文信息。这两个模块都被设计为通过使用非对称卷积来保持低计算复杂度。通过这些模块,我们提出了一个名为“FASSD-Net”的网络,该网络基于轻量级CNN主干。在单个Nvidia GTX 1080Ti上运行,我们的模型分别在Cityscapes和CamVid数据集上以41和80 FPS的速度达到了mIoU的77.5%和69.3%。我们对不同嵌入式系统上的三种FASSD-Net变体的准确性-速度权衡进行了广泛的分析,证明我们的网络的轻型版本可以在低功耗Jetson Xavier NX上运行,以32 FPS的速度达到74%的mIoU,具有全分辨率(1024 × 2048美元)。源代码和预训练模型可在github.com/GibranBenitez/FASSD-Net上获得。
Recent works of real-time semantic segmentation, remove or make use of light decoders from dense deep neural networks to achieve fast inference speed. This strategy helps to achieve real-time performance; however, the accuracy is significantly compromised in comparison to non-real-time methods. In this paper, we introduce two key modules aimed to design a high-performance decoder for real-time semantic segmentation, which also reduces the accuracy gap between real-time and non-real-time networks. The first module, Dilated Asymmetric Pyramidal Fusion (DAPF), is designed to increase the receptive field on the top of the last stage of the encoder, obtaining richer contextual features. The second module, Multi-resolution Dilated Asymmetric (MDA) module, fuses and refines detail and contextual information from multi-scale feature maps coming from early and deeper stages of the network. Both modules are designed to keep a low computational complexity by using asymmetric convolutions. With these modules, we propose a network entitled “FASSD-Net,” which is based on a light-weight CNN backbone. Running on a single Nvidia GTX 1080Ti, our model reaches 77.5% and 69.3% of mIoU, at 41 and 80 FPS on the Cityscapes and CamVid datasets, respectively. We present an extensive analysis of the accuracy-speed tradeoffs of three FASSD-Net variations on different embedded systems, demonstrating that a light version of our network can run on the low-power consumption Jetson Xavier NX, at 32 FPS reaching 74% of mIoU with full resolution ( $1024\times 2048$ ). The source code and pre-trained models are available at github.com/GibranBenitez/FASSD-Net.