Multi-level feature aggregation network for instrument identification of endoscopic images

Multi-level feature aggregation network for instrument identification of endoscopic images
复制标题

用于内窥镜图像器械识别的多级特征聚合网络

DOI:
10.1088/1361-6560/ab8dda
复制
发表时间:
2020-08-21
影响因子:
3.5
通讯作者:
Yang, Jian
Yang, Jian
中科院分区:
工程技术2区
文献类型:
--
作者:
Chu, Yakui;Yang, Xilin;Yang, Jian

文献摘要

被引文献

相似文献

手术器械的识别对于理解手术场景和在内窥镜图像引导手术中提供辅助过程至关重要。这项研究提出了一种新的多级特征聚合深度卷积神经网络(MLFA-Net),用于识别内窥镜图像中的手术器械。首先,在骨干网顶层创建全局特征增强层,将高层语义信息提升到特征流网络,提高对象识别的局部化能力。其次,提出了一种改进的跨通道特征交互路径,增加了同层特征的非线性组合,提高了信息传播的效率。第三,建立多视图特征融合分支,聚合不同视图中同一层次的位置敏感信息,增加特征的信息多样性,增强目标定位能力。通过利用潜在的信息,所提出的多级特征聚合网络可以完成多任务的仪器识别与一个单一的网络。该网络处理三项任务,包括物体检测,对仪器类型进行分类并定位其边界;掩模分割,检测仪器形状;姿态估计,检测仪器部件的关键点。实验在MICCAI 2017内窥镜视觉挑战赛的腹腔镜图像上进行,并利用平均平均精度(AP)和平均召回率(AR)来量化分割和姿态估计结果。对于边界框回归,AP和AR分别为79.1%和63.2%,而掩模分割的AP和AR分别为78.1%和62.1%,姿态估计的AP和AR分别达到67.1%和55.7%。实验表明,该方法有效地提高了内窥镜图像中仪器的识别准确率,优于其他最先进的方法。
Identification of surgical instruments is crucial in understanding surgical scenarios and providing an assistive process in endoscopic image-guided surgery. This study proposes a novel multilevel feature-aggregated deep convolutional neural network (MLFA-Net) for identifying surgical instruments in endoscopic images. First, a global feature augmentation layer is created on the top layer of the backbone to improve the localization ability of object identification by boosting the high-level semantic information to the feature flow network. Second, a modified interaction path of cross-channel features is proposed to increase the nonlinear combination of features in the same level and improve the efficiency of information propagation. Third, a multiview fusion branch of features is built to aggregate the location-sensitive information of the same level in different views, increase the information diversity of features, and enhance the localization ability of objects. By utilizing the latent information, the proposed network of multilevel feature aggregation can accomplish multitask instrument identification with a single network. Three tasks are handled by the proposed network, including object detection, which classifies the type of instrument and locates its border; mask segmentation, which detects the instrument shape; and pose estimation, which detects the keypoint of instrument parts. The experiments are performed on laparoscopic images from MICCAI 2017 Endoscopic Vision Challenge, and the mean average precision (AP) and average recall (AR) are utilized to quantify the segmentation and pose estimation results. For the bounding box regression, the AP and AR are 79.1% and 63.2%, respectively, while the AP and AR of mask segmentation are 78.1% and 62.1%, and the AP and AR of the pose estimation achieve 67.1% and 55.7%, respectively. The experiments demonstrate that our method efficiently improves the recognition accuracy of the instrument in endoscopic images, and outperforms the other state-of-the-art methods.