Real-Time Instrument Segmentation in Robotic Surgery Using Auxiliary Supervised Deep Adversarial Learning

Real-Time Instrument Segmentation in Robotic Surgery Using Auxiliary Supervised Deep Adversarial Learning
复制标题

DOI:
10.1109/lra.2019.2900854
复制
发表时间:
2019-04-01
影响因子:
5.2
通讯作者:
Ren, Hongliang
Ren, Hongliang
中科院分区:
计算机科学2区
文献类型:
--
作者:
Islam, Mobarakol;Atputharuban, Daniel Anojan;Ren, Hongliang

文献摘要

被引文献

相似文献

机器人辅助手术是一项新兴技术,随着机器人技术和成像系统的发展而迅速发展。视觉、触觉和机器人手臂精确运动方面的创新使外科医生能够进行精确的微创手术。机器人手术器械和组织的实时语义分割是机器人辅助手术的关键步骤。手术场景的准确和有效分割不仅有助于识别和跟踪器械,而且还提供了关于正在操作的不同组织和器械的上下文信息。为此,我们开发了一个轻量级的级联卷积神经网络,从商业机器人系统获得的高分辨率视频中分割手术器械。我们提出了一个多分辨率特征融合模块来融合辅助分支和主分支的不同维度和通道的特征图。我们还引入了一种新的方法,结合辅助损失和对抗损失来正则化分割模型。辅助损失有助于模型学习低分辨率特征,而对抗损失通过学习高阶结构信息来改善分割预测。该模型还包括一个轻量级的空间金字塔池单元,以在中间阶段聚合丰富的上下文信息。我们表明,我们的模型优于现有的算法,在高分辨率视频的预测精度和分割时间的像素分割的手术器械。
Robot-assisted surgery is an emerging technology that has undergone rapid growth with the development of robotics and imaging systems. Innovations in vision, haptics, and accurate movements of robot arms have enabled surgeons to perform precise minimally invasive surgeries. Real-time semantic segmentation of the robotic instruments and tissues is a crucial step in robot-assisted surgery. Accurate and efficient segmentation of the surgical scene not only aids in the identification and tracking of instruments but also provides contextual information about the different tissues and instruments being operated with. For this purpose, we have developed a light-weight cascaded convolutional neural network to segment the surgical instruments from high-resolution videos obtained from a commercial robotic system. We propose a multi-resolution feature fusion module to fuse the feature maps of different dimensions and channels from the auxiliary and main branch. We also introduce a novel way of combining auxiliary loss and adversarial loss to regularize the segmentation model. Auxiliary loss helps the model to learn low-resolution features, and adversarial loss improves the segmentation prediction by learning higher order structural information. The model also consists of a lightweight spatial pyramid pooling unit to aggregate rich contextual information in the intermediate stage. We show that our model surpasses existing algorithms for pixelwise segmentation of surgical instruments in both prediction accuracy and segmentation time of high-resolution videos.