Robotic Scene Segmentation with Memory Network for Runtime Surgical Context Inference

Robotic Scene Segmentation with Memory Network for Runtime Surgical Context Inference
复制标题

DOI:
10.1109/iros55552.2023.10342013
复制
发表时间:
2023-08
期刊:
2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Zongyu Li;Ian Reyes;H. Alemzadeh
Zongyu Li;Ian Reyes;H. Alemzadeh
中科院分区:
其他
文献类型:
--
作者:
Zongyu Li;Ian Reyes;H. Alemzadeh

文献摘要

相似文献

手术背景推断最近在机器人辅助手术中引起了极大的关注,因为它可以促进工作流程分析、技能评估和错误检测。然而,运行时上下文推断具有挑战性,因为它需要基于视频数据的分割及时准确地检测手术场景中工具和对象之间的交互。另一方面,现有最先进的视频分割方法通常对不常见的类有偏见,并且无法为分段掩模提供时间一致性。这可能会对上下文推断和关键状态的准确检测产生负面影响。在本研究中,我们提出了使用时空对应网络(STCN)来应对这些挑战的解决方案。 STCN 是一个执行二进制分割并最大限度地减少类别不平衡影响的内存网络。 STCN 中存储体的使用允许利用过去的图像和分割信息,从而确保掩模的一致性。我们使用公开可用的 JIGSAWS 数据集进行的实验表明,STCN 对于难以分割的对象(例如针和线)实现了卓越的分割性能,并且与最先进的技术相比,改进了上下文推理。我们还证明了分段和上下文推断可以在运行时执行,而不会影响性能。
Surgical context inference has recently garnered significant attention in robot-assisted surgery as it can facilitate workflow analysis, skill assessment, and error detection. However, runtime context inference is challenging since it requires timely and accurate detection of the interactions among the tools and objects in the surgical scene based on the segmentation of video data. On the other hand, existing state-of-the-art video segmentation methods are often biased against infrequent classes and fail to provide temporal consistency for segmented masks. This can negatively impact the context inference and accurate detection of critical states. In this study, we propose a solution to these challenges using a Space-Time Correspondence Network (STCN). STCN is a memory network that performs binary segmentation and minimizes the effects of class imbalance. The use of a memory bank in STCN allows for the utilization of past image and segmentation information, thereby ensuring consistency of the masks. Our experiments using the publicly-available JIGSAWS dataset demonstrate that STCN achieves superior segmentation performance for objects that are difficult to segment, such as needle and thread, and improves context inference compared to the state-of-the-art. We also demonstrate that segmentation and context inference can be performed at runtime without compromising performance.