Deep Learning Based Robotic Tool Detection and Articulation Estimation With Spatio-Temporal Layers

Deep Learning Based Robotic Tool Detection and Articulation Estimation With Spatio-Temporal Layers
复制标题

DOI:
10.1109/lra.2019.2917163
复制
发表时间:
2019-07-01
影响因子:
5.2
通讯作者:
Stoyanov, Danail
Stoyanov, Danail
中科院分区:
计算机科学2区
文献类型:
--
作者:
Colleoni, Emanuele;Moccia, Sara;Stoyanov, Danail

文献摘要

被引文献

相似文献

在计算机辅助微创手术中,从腹腔镜图像中检测手术工具接头是一项重要但具有挑战性的任务。光照水平、背景变化以及视野中不同数量的工具,都给算法和模型训练带来了困难。然而,这些挑战可以通过利用腹腔镜视频中的时间信息来避免逐帧处理问题来解决。在这封信中,我们提出了一种用于手术器械联合检测和定位的新型编码器-解码器架构,该架构使用三维卷积层来利用腹腔镜视频的时空特征。在基准和定制数据集上进行测试时,Dice 相似系数中位数为 85.1%,四分位距为 4.6%,凸显了性能优于基于单帧处理的现有技术。除了网络架构的新颖性之外,在训练阶段处理具有不可见背景的图像时,包含时间信息的想法似乎特别有用,这表明用于联合检测的时空特征有助于概括解决方案。
Surgical-tool joint detection from laparoscopic images is an important but challenging task in computer-assisted minimally invasive surgery. Illumination levels, variations in background and the different number of tools in the field of view, all pose difficulties to algorithm and model training. Yet, such challenges could be potentially tackled by exploiting the temporal information in laparoscopic videos to avoid per frame handling of the problem. In this letter, we propose a novel encoder-decoder architecture for surgical instrument joint detection and localization that uses three-dimensional convolutional layers to exploit spatio-temporal features from laparoscopic videos. When tested on benchmark and custom-built datasets, a median Dice similarity coefficient of 85.1% with an interquartile range of 4.6% highlights performance better than the state of the art based on single-frame processing. Alongside novelty of the network architecture, the idea for inclusion of temporal information appears to be particularly useful when processing images with unseen backgrounds during the training phase, which indicates that spatio-temporal features for joint detection help to generalize the solution.