Weakly supervised convolutional LSTM approach for tool tracking in laparoscopic videos

Weakly supervised convolutional LSTM approach for tool tracking in laparoscopic videos
复制标题

DOI:
10.1007/s11548-019-01958-6
复制
发表时间:
2019-06-01
影响因子:
3
通讯作者:
Padoy, Nicolas
Padoy, Nicolas
中科院分区:
工程技术3区
文献类型:
--
作者:
Nwoye, Chinedu Innocent;Mutter, Didier;Padoy, Nicolas

文献摘要

被引文献

相似文献

手术工具的实时跟踪是未来智能手术室的核心组成部分,因为它对于分析和理解手术活动具有重要意义。当前用于视频中的手术工具跟踪的方法需要在工具的空间位置被手动注释的数据上进行训练。生成这样的训练数据是困难且耗时的。相反,我们建议只使用二进制存在注释来训练腹腔镜videos.MethodsThe所提出的方法是由CNN +卷积LSTM(ConvLSTM)神经网络训练端到端,但弱监督工具二进制存在标签。我们使用ConvLSTM对手术工具运动中的时间依赖性进行建模,并利用其时空能力来平滑定位热图(Lh-maps)中的类峰激活。ResultsWe在CNN模型之上构建基线跟踪器,并证明我们基于ConvLSTM的方法在工具存在检测,空间定位和运动跟踪方面优于基线超过5.0%,13.9%和12.6%。结论在本文中,我们证明了二进制存在标签足以使用我们提出的方法训练深度学习跟踪模型。我们还表明,ConvLSTM可以利用手术视频中连续图像帧的时空相干性来改善工具存在检测、空间定位和运动跟踪。
PurposeReal-time surgical tool tracking is a core component of the future intelligent operating room (OR), because it is highly instrumental to analyze and understand the surgical activities. Current methods for surgical tool tracking in videos need to be trained on data in which the spatial positions of the tools are manually annotated. Generating such training data is difficult and time-consuming. Instead, we propose to use solely binary presence annotations to train a tool tracker for laparoscopic videos.MethodsThe proposed approach is composed of a CNN + Convolutional LSTM (ConvLSTM) neural network trained end to end, but weakly supervised on tool binary presence labels only. We use the ConvLSTM to model the temporal dependencies in the motion of the surgical tools and leverage its spatiotemporal ability to smooth the class peak activations in the localization heat maps (Lh-maps).ResultsWe build a baseline tracker on top of the CNN model and demonstrate that our approach based on the ConvLSTM outperforms the baseline in tool presence detection, spatial localization, and motion tracking by over 5.0%, 13.9%, and 12.6%, respectively.ConclusionsIn this paper, we demonstrate that binary presence labels are sufficient for training a deep learning tracking model using our proposed method. We also show that the ConvLSTM can leverage the spatiotemporal coherence of consecutive image frames across a surgical video to improve tool presence detection, spatial localization, and motion tracking.