Keep Your Eye on the Best: Contrastive Regression Transformer for Skill Assessment in Robotic Surgery

Keep Your Eye on the Best: Contrastive Regression Transformer for Skill Assessment in Robotic Surgery
复制标题

DOI:
10.1109/lra.2023.3242466
复制
发表时间:
2023-03-01
影响因子:
5.2
通讯作者:
Mazomenos, Evangelos
Mazomenos, Evangelos
中科院分区:
计算机科学2区
文献类型:
--
作者:
Anastasiou, Dimitrios;Jin, Yueming;Mazomenos, Evangelos

文献摘要

被引文献

相似文献

这封信提出了一种新的基于视频的,对比回归架构,对比变形器,在机器人辅助手术的自动手术技能评估。所提出的框架的结构,以捕捉手术性能的差异,测试视频和参考视频,代表最佳的手术执行。特征提取器将在帧级上用手势标签监督的空间分量(ResNet-18)和时间分量(TCN)组合,生成测试和参考视频的时空特征矩阵。然后,这些被馈送到具有多头注意力的动作感知Transformer中,该动作感知transformer在帧级产生视频间对比特征,表示两个视频之间的技能相似性/偏差。次优性能的时刻可以被识别并且在所获得的特征向量中在时间上被定位,其最终用于回归手动分配的技能分数。在JIGSAWS数据集上验证,Contra-Transformer实现了具有竞争力的性能(斯皮尔曼0.65-0.89),所有任务和验证设置的归一化平均绝对误差在5.8%-13.4%之间。
This letter proposes a novel video-based, contrastive regression architecture, Contra-Sformer, for automated surgical skill assessment in robot-assisted surgery. The proposed framework is structured to capture the differences in the surgical performance, between a test video and a reference video which represents optimal surgical execution. A feature extractor combining a spatial component (ResNet-18), supervised on frame-level with gesture labels, and a temporal component (TCN), generates spatio-temporal feature matrices of the test and reference videos. These are then fed into an action-aware Transformer with multi-head attention that produces inter-video contrastive features at frame level, representative of the skill similarity/deviation between the two videos. Moments of sub-optimal performance can be identified and temporally localized in the obtained feature vectors, which are ultimately used to regress the manually assigned skill scores. Validated on the JIGSAWS dataset, Contra-Sformer achieves competitive performance (Spearman 0.65-0.89), with a normalized mean absolute error between 5.8%-13.4% on all tasks and across validation setups.