Recognition and Prediction of Surgical Gestures and Trajectories Using Transformer Models in Robot-Assisted Surgery

Recognition and Prediction of Surgical Gestures and Trajectories Using Transformer Models in Robot-Assisted Surgery
复制标题

DOI:
10.1109/iros47612.2022.9981611
复制
发表时间:
2022-10
期刊:
2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS)
影响因子:
--
通讯作者:
Chang Shi;Y. Zheng;A. M. Fey
Chang Shi;Y. Zheng;A. M. Fey
中科院分区:
其他
文献类型:
--
作者:
Chang Shi;Y. Zheng;A. M. Fey

文献摘要

相似文献

手术活动识别和预测可以帮助在许多机器人辅助手术(RAS)应用中提供重要的上下文,例如,手术进度监测和估计、手术技能评估以及远程操作期间的共享控制策略。Transformer模型最初是为自然语言处理(NLP)开发的,用于对单词序列进行建模,很快该方法在一般序列建模任务中得到了普及。在本文中,我们提出了一个新的使用Transformer模型的三个任务:手势识别,手势预测和轨迹预测在RAS。我们修改原来的Transformer架构,能够生成当前的姿态序列,未来的姿态序列,未来的轨迹序列估计,只使用当前的运动学数据的手术机器人末端执行器。我们在JHU-ISI手势和技能评估工作集(JIGSAWS)上评估我们提出的模型,并使用Leave-One-User-Out(LOUO)交叉验证来确保我们结果的普遍性。我们的模型实现了高达89.3%的手势识别准确率,84.6%的手势预测准确率(提前1秒)和2.71mm的轨迹预测误差(提前1秒)。我们的模型是可比的,并能够超越国家的最先进的方法,同时只使用运动学数据通道。这种方法可以实现近实时的手术活动识别和预测。
Surgical activity recognition and prediction can help provide important context in many Robot-Assisted Surgery (RAS) applications, for example, surgical progress monitoring and estimation, surgical skill evaluation, and shared control strategies during teleoperation. Transformer models were first developed for Natural Language Processing (NLP) to model word sequences and soon the method gained popularity for general sequence modeling tasks. In this paper, we propose the novel use of a Transformer model for three tasks: gesture recognition, gesture prediction, and trajectory prediction during RAS. We modify the original Transformer architecture to be able to generate the current gesture sequence, future gesture sequence, and future trajectory sequence estimations using only the current kinematic data of the surgical robot end-effectors. We evaluate our proposed models on the JHU-ISI Gesture and Skill Assessment Working Set (JIGSAWS) and use Leave-One-User-Out (LOUO) cross validation to ensure generalizability of our results. Our models achieve up to 89.3% gesture recognition accuracy, 84.6% gesture prediction accuracy (1 second ahead) and 2.71mm trajectory prediction error (1 second ahead). Our models are comparable to and able to outperform state-of-the-art methods while using only the kinematic data channel. This approach can enable near-real time surgical activity recognition and prediction.