Lightweight Underwater Object Detection Based on YOLO v4 and Multi-Scale Attentional Feature Fusion

Lightweight Underwater Object Detection Based on YOLO v4 and Multi-Scale Attentional Feature Fusion
复制标题

DOI:
10.3390/rs13224706
复制
发表时间:
2021-11-01
期刊:
影响因子:
5
通讯作者:
Wei, Quanmiao
Wei, Quanmiao
中科院分区:
工程技术2区
文献类型:
--
作者:
Zhang, Minghua;Xu, Shubo;Wei, Quanmiao

文献摘要

被引文献

相似文献

水下目标检测是计算机视觉中一个具有挑战性和吸引力的课题。虽然目标检测技术在一般数据集上取得了较好的性能,但复杂水下环境中的低可见度和颜色偏差问题导致图像质量普遍较差;此外,小目标和目标聚集问题导致可提取的信息较少,难以达到令人满意的效果。在以往基于深度学习的水下目标检测研究中,大多数研究主要集中在利用大型网络提高检测精度,而海洋水下轻量级目标检测问题很少受到关注,导致模型规模大,检测速度慢,因此海洋环境下目标检测技术的应用需要更好的实时性和轻量级性能。针对这一问题,提出了一种基于MobileNet v2、You Only Look Once(YOLO)v4算法和注意力特征融合的轻量级水下目标检测方法,实现了海洋环境下目标检测准确性和快速性的和谐平衡。在我们的工作中,提出了MobileNet v2和深度可分离卷积的组合,以减少模型参数的数量和模型的大小。改进的注意力特征融合(AFFM)模块旨在更好地融合语义和尺度不一致的特征,并提高准确性。实验结果表明,该方法在PASCAL VOC数据集和半咸水数据集上的平均精度分别达到81.67%和92.65%,在半咸水数据集上的处理速度达到44.22帧/秒(FPS).此外,模型参数的数量和模型大小分别压缩到YOLO v4的16.76%和19.53%,这实现了水下目标检测的时间和精度之间的良好折衷。
A challenging and attractive task in computer vision is underwater object detection. Although object detection techniques have achieved good performance in general datasets, problems of low visibility and color bias in the complex underwater environment have led to generally poor image quality; besides this, problems with small targets and target aggregation have led to less extractable information, which makes it difficult to achieve satisfactory results. In past research of underwater object detection based on deep learning, most studies have mainly focused on improving detection accuracy by using large networks; the problem of marine underwater lightweight object detection has rarely gotten attention, which has resulted in a large model size and slow detection speed; as such the application of object detection technologies under marine environments needs better real-time and lightweight performance. In view of this, a lightweight underwater object detection method based on the MobileNet v2, You Only Look Once (YOLO) v4 algorithm and attentional feature fusion has been proposed to address this problem, to produce a harmonious balance between accuracy and speediness for target detection in marine environments. In our work, a combination of MobileNet v2 and depth-wise separable convolution is proposed to reduce the number of model parameters and the size of the model. The Modified Attentional Feature Fusion (AFFM) module aims to better fuse semantic and scale-inconsistent features and to improve accuracy. Experiments indicate that the proposed method obtained a mean average precision (mAP) of 81.67% and 92.65% on the PASCAL VOC dataset and the brackish dataset, respectively, and reached a processing speed of 44.22 frame per second (FPS) on the brackish dataset. Moreover, the number of model parameters and the model size were compressed to 16.76% and 19.53% of YOLO v4, respectively, which achieved a good tradeoff between time and accuracy for underwater object detection.