LMFFNet: A Well-Balanced Lightweight Network for Fast and Accurate Semantic Segmentation

LMFFNet: A Well-Balanced Lightweight Network for Fast and Accurate Semantic Segmentation
复制标题

DOI:
10.1109/tnnls.2022.3176493
复制
发表时间:
2022-05
影响因子:
10.4
通讯作者:
Min Shi;Jialin Shen;Qingming Yi;Jian Weng;Zunkai Huang;Aiwen Luo;Yicong Zhou
Min Shi;Jialin Shen;Qingming Yi;Jian Weng;Zunkai Huang;Aiwen Luo;Yicong Zhou
中科院分区:
计算机科学1区
文献类型:
--
作者:
Min Shi;Jialin Shen;Qingming Yi;Jian Weng;Zunkai Huang;Aiwen Luo;Yicong Zhou

文献摘要

相似文献

实时语义分割广泛应用于自动驾驶和机器人领域。大多数先前的网络基于涉及大规模计算的复杂模型而实现了很高的精度。现有的轻量级网络通常通过牺牲分割精度来减少参数大小。平衡实时语义分割的参数和准确性至关重要。在本文中,我们提出了一种轻量级多尺度特征融合网络(LMFFNet),主要由三种类型的组件组成:分割提取合并瓶颈(SEM-B)块、特征融合模块(FFM)和多尺度注意解码器(MAD),其中SEM-B块用更少的参数提取足够的特征。 FFM融合多尺度语义特征,有效提高分割精度,MAD通过注意力机制很好地恢复了输入图像的细节。在没有预训练的情况下,LMFFNet-3-8 使用 RTX 3090 GPU 在 118.9 帧/秒的速度下以 140 万个参数实现了 75.1% 的并集平均交集 (mIoU)。在 CamVid、KITTI 和 WildDash2 等其他三个数据集上的各种分辨率上进行了更多实验。实验验证了所提出的 LMFFNet 模型在实时任务的分割精度和推理速度之间做出了不错的权衡。源代码可在 https://github.com/Greak-1124/LMFFNet 上公开获取。
Real-time semantic segmentation is widely used in autonomous driving and robotics. Most previous networks achieved great accuracy based on a complicated model involving mass computing. The existing lightweight networks generally reduce the parameter sizes by sacrificing the segmentation accuracy. It is critical to balance the parameters and accuracy for real-time semantic segmentation. In this article, we propose a lightweight multiscale-feature-fusion network (LMFFNet) mainly composed of three types of components: split-extract-merge bottleneck (SEM-B) block, feature fusion module (FFM), and multiscale attention decoder (MAD), where the SEM-B block extracts sufficient features with fewer parameters. FFMs fuse multiscale semantic features to effectively improve the segmentation accuracy and the MAD well recovers the details of the input images through the attention mechanism. Without pretraining, LMFFNet-3-8 achieves 75.1% mean intersection over union (mIoU) with 1.4 M parameters at 118.9 frames/s using RTX 3090 GPU. More experiments are investigated extensively on various resolutions on other three datasets of CamVid, KITTI, and WildDash2. The experiments verify that the proposed LMFFNet model makes a decent tradeoff between segmentation accuracy and inference speed for real-time tasks. The source code is publicly available at https://github.com/Greak-1124/LMFFNet.