SV-RCNet: Workflow Recognition From Surgical Videos Using Recurrent Convolutional Network

SV-RCNet: Workflow Recognition From Surgical Videos Using Recurrent Convolutional Network
复制标题

DOI:
10.1109/tmi.2017.2787657
复制
发表时间:
2018-05-01
影响因子:
10.6
通讯作者:
Heng, Pheng-Ann
Heng, Pheng-Ann
中科院分区:
工程技术1区
文献类型:
--
作者:
Jin, Yueming;Dou, Qi;Heng, Pheng-Ann

文献摘要

被引文献

相似文献

我们提出了一种基于新型递归卷积网络(SV-RCNet)的手术视频分析,特别是用于从在线手术视频中自动识别工作流,这是开发上下文感知计算机辅助干预系统的关键组成部分。与以往分别利用视觉和时间信息的方法不同,本文提出的SV-RCNet将卷积神经网络(CNN)和递归神经网络(RNN)无缝集成,形成一种新的递归卷积架构,以充分利用从手术视频中学习到的视觉和时间特征的互补信息。我们以端到端方式有效地训练SV-RCNet,从而在学习过程中共同优化视觉表征和序列动态。为了产生更具判别性的时空特征,我们利用深度残差网络(ResNet)和长短期记忆网络(LSTM)分别提取视觉特征和时间依赖性,并将它们整合到SV-RCNet中。此外,基于SV-RCNet的相变敏感预测,我们提出了一种简单而有效的推理方案,即利用手术视频的自然特征的先验知识推理(PKI)。这种策略进一步提高了结果的一致性,极大地提高了识别性能。利用MICCAI 2016计算机辅助干预建模与监测工作流挑战数据集和Cholec80数据集进行了大量实验,以验证SV-RCNet。我们的方法不仅在这两个数据集上实现了卓越的性能,而且在很大程度上优于最先进的方法。
We propose an analysis of surgical videos that is based on a novel recurrent convolutional network (SV-RCNet), specifically for automatic workflow recognition from surgical videos online, which is a key component for developing the context-aware computer-assisted intervention systems. Different from previous methods which harness visual and temporal information separately, the proposed SV-RCNet seamlessly integrates a convolutional neural network (CNN) and a recurrent neural network (RNN) to forma novel recurrent convolutional architecture in order to take full advantages of the complementary information of visual and temporal features learned from surgical videos. We effectively train the SV-RCNet in an end-to-end manner so that the visual representations and sequential dynamics can be jointly optimized in the learning process. In order to produce more discriminative spatio-temporal features, we exploit a deep residual network (ResNet) and a long short term memory (LSTM) network, to extract visual features and temporal dependencies, respectively, and integrate them into the SV-RCNet. Moreover, based on the phase transition-sensitive predictions from the SV-RCNet, we propose a simple yet effective inference scheme, namely the prior knowledge inference (PKI), by leveraging the natural characteristic of surgical video. Such a strategy further improves the consistency of results and largely boosts the recognition performance. Extensive experiments have been conducted with the MICCAI 2016 Modeling and Monitoring of Computer Assisted Interventions Workflow Challenge dataset and Cholec80 dataset to validate SV-RCNet. Our approach not only achieves superior performance on these two datasets but also outperforms the state-of-the-art methods by a significant margin.