Open surgery tool classification and hand utilization using a multi-camera system.

Open surgery tool classification and hand utilization using a multi-camera system.
复制标题

DOI:
10.1007/s11548-022-02691-3
复制
发表时间:
2022-08
影响因子:
3
通讯作者:
--
中科院分区:
工程技术3区
文献类型:
--
作者:

文献摘要

被引文献

相似文献

这项工作的目标是使用多摄像头视频对开放手术工具进行分类,并确定每只手拿着哪个工具。多摄像头系统有助于防止闭塞在开放手术视频数据。此外,结合多个视图,如覆盖整个手术领域的俯视图摄像机和聚焦手部运动和解剖的特写摄像机,可以提供更全面的手术工作流程视图。然而,多相机数据融合带来了一个新的挑战:一个工具可能在一个相机上可见,而在另一个相机上不可见。因此,我们将全球基础真相定义为正在使用的工具,而不考虑其可见性。因此,当系统对视频中可见的变化做出快速响应时,应该长时间记住图像之外的工具。参与者(n=48)进行了模拟开放肠修复。使用了一个俯视图和一个特写摄像机。使用YOLOv5进行工具和手的检测。高频LSTM和低频LSTM分别用于空间、时间和多相机集成,前者具有1秒窗口,帧率为30帧/秒,后者具有40秒窗口,帧率为3帧/秒。六个系统的精度和F1分别为:俯视图(0.88/0.88)、特写(0.81、0.83)、双相机(0.9/0.9)、高fps LSTM(0.92/0.93)、低fps LSTM(0.9/0.91)和我们最终的多相机分类器(0.93/0.94)。通过将多相机阵列的高fps和低fps相结合,提高了对全局地面真值的分类能力。
The goal of this work is to use multi-camera video to classify open surgery tools as well as identify which tool is held in each hand. Multi-camera systems help prevent occlusions in open surgery video data. Furthermore, combining multiple views such as a Top-view camera covering the full operative field and a Close-up camera focusing on hand motion and anatomy, may provide a more comprehensive view of the surgical workflow. However, multi-camera data fusion poses a new challenge: a tool may be visible in one camera and not the other. Thus, we defined the global ground truth as the tools being used regardless their visibility. Therefore, tools that are out of the image should be remembered for extensive periods of time while the system responds quickly to changes visible in the video. Participants (n=48) performed a simulated open bowel repair. A Top-view and a Close-up cameras were used. YOLOv5 was used for tool and hand detection. A high frequency LSTM with a 1 second window at 30 frames per second (fps) and a low frequency LSTM with a 40 second window at 3 fps were used for spatial, temporal, and multi-camera integration. The accuracy and F1 of the six systems were: Top-view (0.88/0.88), Close-up (0.81,0.83), both cameras (0.9/0.9), high fps LSTM (0.92/0.93), low fps LSTM (0.9/0.91), and our final architecture the Multi-camera classifier(0.93/0.94). By combining a system with a high fps and a low fps from the multiple camera array we improved the classification abilities of the global ground truth.