Multimodal Transformers for Real-Time Surgical Activity Prediction

Multimodal Transformers for Real-Time Surgical Activity Prediction
复制标题

DOI:
10.48550/arxiv.2403.06705
复制
发表时间:
2024-03
期刊:
ArXiv
影响因子:
--
通讯作者:
Keshara Weerasinghe;Seyed Hamid Reza Roodabeh;Kay Hutchinson;H. Alemzadeh
Keshara Weerasinghe;Seyed Hamid Reza Roodabeh;Kay Hutchinson;H. Alemzadeh
中科院分区:
其他
文献类型:
--
作者:
Keshara Weerasinghe;Seyed Hamid Reza Roodabeh;Kay Hutchinson;H. Alemzadeh

文献摘要

相似文献

实时识别和预测手术活动是提高机器人辅助手术安全性和自主性的基础。本文提出了一种基于短段运动和视频数据的手术手势和轨迹实时识别和预测的多模式转换器结构。我们进行了一项消融研究,以评估融合不同输入模式及其表示对手势识别和预测性能的影响。我们使用JHU-ISI手势和技能评估工作集(JIGSAWS)数据集对建议的体系结构进行端到端评估。通过将运动学特征与空间和上下文视频特征进行有效融合,我们的模型在手势预测方面的准确率超过了最新的SOTA(89.5\%)。它依靠计算效率高的模型,实现了处理1秒输入窗口的1.1-1.3ms的实时性能。
Real-time recognition and prediction of surgical activities are fundamental to advancing safety and autonomy in robot-assisted surgery. This paper presents a multimodal transformer architecture for real-time recognition and prediction of surgical gestures and trajectories based on short segments of kinematic and video data. We conduct an ablation study to evaluate the impact of fusing different input modalities and their representations on gesture recognition and prediction performance. We perform an end-to-end assessment of the proposed architecture using the JHU-ISI Gesture and Skill Assessment Working Set (JIGSAWS) dataset. Our model outperforms the state-of-the-art (SOTA) with 89.5\% accuracy for gesture prediction through effective fusion of kinematic features with spatial and contextual video features. It achieves the real-time performance of 1.1-1.3ms for processing a 1-second input window by relying on a computationally efficient model.