A Real-Time Apple Targets Detection Method for Picking Robot Based on Improved YOLOv5

A Real-Time Apple Targets Detection Method for Picking Robot Based on Improved YOLOv5
复制标题

DOI:
10.3390/rs13091619
复制
发表时间:
2021-05-01
期刊:
影响因子:
5
通讯作者:
Yang, Fuzeng
Yang, Fuzeng
中科院分区:
工程技术2区
文献类型:
--
作者:
Yan, Bin;Fan, Pan;Yang, Fuzeng

文献摘要

被引文献

相似文献

苹果目标识别算法是摘苹果机器人的核心技术之一。然而,现有的大多数苹果检测算法无法区分被树枝遮挡的苹果和被其他苹果遮挡的苹果。如果将该算法直接应用到采摘机器人上,机器人的苹果、抓取端执行器和机械采摘臂很可能会损坏。基于这一实际问题,为了自动识别苹果树图像中可抓取和不可抓取的苹果,提出了一种基于改进YOLOv5s的摘苹果机器人轻量目标检测方法。首先,将瓶颈csp模块改进设计为瓶颈csp -2模块,用于替代原有YOLOv5s网络骨干架构中的瓶颈csp模块。其次,将视觉注意机制网络中的SE模块插入到改进后的骨干网中;第三,改进了原YOLOv5s网络中中等大小目标检测层输入的特征映射的键合融合模式。最后,对原网络初始锚盒尺寸进行改进。实验结果表明,本文提出的改进网络模型可以有效识别未被树叶遮挡或仅被树叶遮挡的可抓取苹果,以及被树枝遮挡或被其他水果遮挡的不可抓取苹果。其中,识别查全率为91.48%,查准率为83.83%,mAP为86.75%,F1为87.49%。平均识别时间为0.015 s /张。与原来的YOLOv5s、YOLOv3、YOLOv4和efficientdot - d0模型相比,改进后的YOLOv5s模型的mAP分别提高了5.05%、14.95%、4.74%和6.75%,压缩后的模型大小分别提高了9.29%、94.6%、94.8%和15.3%。改进后的YOLOv5s模型的平均单幅图像识别速度分别是efficientet - d0、YOLOv4和YOLOv3和model的2.53倍、1.13倍和3.53倍。该方法可为苹果采摘机器人实时准确检测多个水果目标提供技术支持。
The apple target recognition algorithm is one of the core technologies of the apple picking robot. However, most of the existing apple detection algorithms cannot distinguish between the apples that are occluded by tree branches and occluded by other apples. The apples, grasping end-effector and mechanical picking arm of the robot are very likely to be damaged if the algorithm is directly applied to the picking robot. Based on this practical problem, in order to automatically recognize the graspable and ungraspable apples in an apple tree image, a light-weight apple targets detection method was proposed for picking robot using improved YOLOv5s. Firstly, BottleneckCSP module was improved designed to BottleneckCSP-2 module which was used to replace the BottleneckCSP module in backbone architecture of original YOLOv5s network. Secondly, SE module, which belonged to the visual attention mechanism network, was inserted to the proposed improved backbone network. Thirdly, the bonding fusion mode of feature maps, which were inputs to the target detection layer of medium size in the original YOLOv5s network, were improved. Finally, the initial anchor box size of the original network was improved. The experimental results indicated that the graspable apples, which were unoccluded or only occluded by tree leaves, and the ungraspable apples, which were occluded by tree branches or occluded by other fruits, could be identified effectively using the proposed improved network model in this study. Specifically, the recognition recall, precision, mAP and F1 were 91.48%, 83.83%, 86.75% and 87.49%, respectively. The average recognition time was 0.015 s per image. Contrasted with original YOLOv5s, YOLOv3, YOLOv4 and EfficientDet-D0 model, the mAP of the proposed improved YOLOv5s model increased by 5.05%, 14.95%, 4.74% and 6.75% respectively, the size of the model compressed by 9.29%, 94.6%, 94.8% and 15.3% respectively. The average recognition speeds per image of the proposed improved YOLOv5s model were 2.53, 1.13 and 3.53 times of EfficientDet-D0, YOLOv4 and YOLOv3 and model, respectively. The proposed method can provide technical support for the real-time accurate detection of multiple fruit targets for the apple picking robot.