Multi-task recurrent convolutional network with correlation loss for surgical video analysis

Multi-task recurrent convolutional network with correlation loss for surgical video analysis
复制标题

用于手术视频分析的具有相关性损失的多任务循环卷积网络

DOI:
10.1016/j.media.2019.101572
复制
发表时间:
2020-01-01
影响因子:
10.9
通讯作者:
Heng, Pheng-Ann
Heng, Pheng-Ann
中科院分区:
工程技术1区
文献类型:
--
作者:
Jin, Yueming;Li, Huaxia;Heng, Pheng-Ann

文献摘要

被引文献

相似文献

手术工具存在检测和手术阶段识别是手术视频分析中两个基本但具有挑战性的任务,也是现代手术室各种应用中非常重要的组成部分。虽然这两个分析任务在临床实践中高度相关,因为手术过程通常是明确定义的,但大多数以前的方法单独处理它们,而没有充分利用它们的相关性。在本文中,我们提出了一种新的方法,通过开发具有相关性损失的多任务递归卷积网络(MTRCNet-CL)来利用它们的相关性,同时提高两个任务的性能。具体来说,我们提出的MTRCNet-CL模型具有端到端架构,其中有两个分支,它们共享早期的特征编码器以提取一般的视觉特征,同时保持针对特定任务的相应高层。鉴于时间信息对于相位识别至关重要,探索长短期记忆(LSTM)来对相位识别分支中的顺序依赖性进行建模。更重要的是,一种新的和有效的相关性损失的设计模型之间的相关性工具的存在和每个视频帧的相位识别,通过最大限度地减少从两个分支的预测的分歧。MTRCNet-CL方法可以在很大程度上鼓励两个任务之间的交互,从而可以为彼此带来好处。在大型手术视频数据集(Cholec 80)上的广泛实验证明了我们提出的方法的出色性能,始终超过最先进的方法,例如,在工具存在检测中,mAP为89.1%vs.81.0%;在相位识别中,F1评分为87.4%vs.84.5%。(C)2019 Elsevier B. V.版权所有。
Surgical tool presence detection and surgical phase recognition are two fundamental yet challenging tasks in surgical video analysis as well as very essential components in various applications in modern operating rooms. While these two analysis tasks are highly correlated in clinical practice as the surgical process is typically well-defined, most previous methods tackled them separately, without making full use of their relatedness. In this paper, we present a novel method by developing a multi-task recurrent convolutional network with correlation loss (MTRCNet-CL) to exploit their relatedness to simultaneously boost the performance of both tasks. Specifically, our proposed MTRCNet-CL model has an end-to-end architecture with two branches, which share earlier feature encoders to extract general visual features while holding respective higher layers targeting for specific tasks. Given that temporal information is crucial for phase recognition, long-short term memory (LSTM) is explored to model the sequential dependencies in the phase recognition branch. More importantly, a novel and effective correlation loss is designed to model the relatedness between tool presence and phase identification of each video frame, by minimizing the divergence of predictions from the two branches. Mutually leveraging both low-level feature sharing and high-level prediction correlating, our MTRCNet-CL method can encourage the interactions between the two tasks to a large extent, and hence can bring about benefits to each other. Extensive experiments on a large surgical video dataset (Cholec80) demonstrate outstanding performance of our proposed method, consistently exceeding the state-of-the-art methods by a large margin, e.g., 89.1%v.s.81.0% for the mAP in tool presence detection and 87.4%v.s.84.5% for F1 score in phase recognition. (C) 2019 Elsevier B.V. All rights reserved.