Evaluating the Task Generalization of Temporal Convolutional Networks for Surgical Gesture and Motion Recognition Using Kinematic Data

Evaluating the Task Generalization of Temporal Convolutional Networks for Surgical Gesture and Motion Recognition Using Kinematic Data
复制标题

DOI:
10.1109/lra.2023.3292581
复制
发表时间:
2023-06
影响因子:
5.2
通讯作者:
Kay Hutchinson;Ian Reyes;Zongyu Li;H. Alemzadeh
Kay Hutchinson;Ian Reyes;Zongyu Li;H. Alemzadeh
中科院分区:
计算机科学2区
文献类型:
--
作者:
Kay Hutchinson;Ian Reyes;Zongyu Li;H. Alemzadeh

文献摘要

被引文献

相似文献

细粒度的活动识别可以对机器人辅助手术中的技能评估、自主性和错误检测程序进行可解释的分析。然而,现有的识别模型受到有限的可用性与运动学和视频数据的注释数据集和无法推广到看不见的主题和任务。来自手术机器人的运动学数据对于安全监控和自主性特别重要,因为它不受遮挡和透镜污染等常见相机问题的影响。我们利用来自总共28名受试者的6个干实验室手术任务的聚合数据集,在手势和运动基元(MP)水平上训练活动识别模型,并仅使用运动学数据训练单独的机器人手臂。使用LOUO(Leave-One-User-Out)和我们提出的LOTO(Leave-One-Task-Out)交叉验证方法对模型进行评估,以评估它们分别推广到看不见的用户和任务的能力。手势识别模型比MP识别模型实现更高的准确性和编辑分数。但是,使用MP可以训练模型,这些模型可以更好地推广到看不见的任务。此外,通过训练用于左机器人臂和右机器人臂的单独模型,可以实现更高的MP识别精度。对于任务泛化,MP识别模型如果在相似的任务和/或来自相同数据集的任务上训练,则表现最好。
Fine-grained activity recognition enables explainable analysis of procedures for skill assessment, autonomy, and error detection in robot-assisted surgery. However, existing recognition models suffer from the limited availability of annotated datasets with both kinematic and video data and an inability to generalize to unseen subjects and tasks. Kinematic data from the surgical robot is particularly critical for safety monitoring and autonomy, as it is unaffected by common camera issues such as occlusions and lens contamination. We leverage an aggregated dataset of six dry-lab surgical tasks from a total of 28 subjects to train activity recognition models at the gesture and motion primitive (MP) levels and for separate robotic arms using only kinematic data. The models are evaluated using the LOUO (Leave-One-User-Out) and our proposed LOTO (Leave-One-Task-Out) cross validation methods to assess their ability to generalize to unseen users and tasks respectively. Gesture recognition models achieve higher accuracies and edit scores than MP recognition models. But, using MPs enables the training of models that can generalize better to unseen tasks. Also, higher MP recognition accuracy can be achieved by training separate models for the left and right robot arms. For task-generalization, MP recognition models perform best if trained on similar tasks and/or tasks from the same dataset.