Evaluation of Deep Learning Models for Identifying Surgical Actions and Measuring Performance

Evaluation of Deep Learning Models for Identifying Surgical Actions and Measuring Performance
复制标题

DOI:
10.1001/jamanetworkopen.2020.1664
复制
发表时间:
2020-03-30
期刊:
影响因子:
13.8
通讯作者:
Rudzicz, Frank
Rudzicz, Frank
中科院分区:
医学1区
文献类型:
--
作者:
Khalid, Shuja;Goldenberg, Mitchell;Rudzicz, Frank

文献摘要

被引文献

相似文献

深度机器学习模型能否用于评估重要的手术特征,例如手术类型和手术性能?结果在这项由8名外科医生进行的103个桌面外科手术视频剪辑的质量改进研究中,包括3个手术操作的4到5个试验,深度机器学习在检测手术动作方面获得了0.97的平均精度和0.98的平均召回率,在估计手术技能水平方面获得了0.77的平均精度和0.78的平均召回率。运营商意义在这项研究中,通过深度机器学习自动处理短手术视频片段准确地识别和评估了手术性能。这项质量改进研究提出了一个机器学习框架,用于评估手术视频片段,方法是根据正在执行的手术步骤和外科医生的能力水平对其进行分类。重要性在手术室评估外科医生时,经验丰富的医生必须依靠现场或录制的视频来评估外科医生的技术表现,这种方法容易产生主观性和错误。由于每天进行大量的外科手术,不可能对每一个手术进行审查;因此,大量丢失了宝贵的性能数据,否则这些数据将有助于提高手术安全性。目的评价一种基于手术步骤和外科医生能力水平对手术视频片段进行分类的评估框架。设计、设置和参与者这项质量改进研究评估了来自约翰霍普金斯大学直观手术手势和技能评估工作集的8名不同级别的外科医生进行打结、打结和穿针的103个视频剪辑。数据收集于2015年之前,数据分析于2019年3月至7月进行。训练深度学习模型以估计分类输出,如性能水平(即新手、中级和专家)和手术操作(即打结、缝合和穿针)。这些模型的有效性通过精确度、召回率和模型准确度来衡量。结果所提供的架构实现了手术动作和性能计算任务的准确性,仅使用视频输入。包埋表示法的平均(均方根误差[RMSE])精密度为1.00(0),打结为0.99(0.01),穿针为0.91(0.11),因此平均(RMSE)精密度为0.97(0.01)。它的平均(RMSE)回忆是0.94(0.08),对于打结,1.00(0),对于穿针,0.99(0.01),导致平均(RMSE)回忆为0.98(0.01)。它还估计了技术技能客观结构化评估全球评级量表类别的分数,新手水平的平均(RMSE)精度为0.85(0.09),中级水平为0.67(0.07),专家水平为0.79(0.12),平均(RMSE)精度为0.77(0.04)。它的平均(RMSE)回忆为0.85(0.05)为新手水平,0.69(0.14)为中级水平,0.80(0.13)为专家水平,导致平均(RMSE)回忆为0.78(0.03)。结论和相关性提出的模型和附带的结果表明,深度机器学习可以识别手术视频片段中的关联。这些是为外科医生创建反馈机制的第一步,这将使他们能够从经验中学习并改进他们的技能。
Question Can deep machine learning models be used to assess important surgical characteristics, such as the type of procedure and surgical performance? Findings In this quality improvement study of 103 video clips of table-top surgical procedures, performed by 8 surgeons and including 4 to 5 trials of 3 surgical actions, deep machine learning obtained a mean precision of 0.97 and a mean recall of 0.98 in detecting surgical actions and a mean precision of 0.77 and a mean recall of 0.78 in estimating the surgical skill level of operators. Meaning In this study, automatic processing of short surgical video clips by deep machine learning accurately identified and assessed surgical performance.This quality improvement study presents a machine learning framework for assessing surgical video clips by categorizing them based on the surgical step being performed and the level of the surgeon's competence.Importance When evaluating surgeons in the operating room, experienced physicians must rely on live or recorded video to assess the surgeon's technical performance, an approach prone to subjectivity and error. Owing to the large number of surgical procedures performed daily, it is infeasible to review every procedure; therefore, there is a tremendous loss of invaluable performance data that would otherwise be useful for improving surgical safety. Objective To evaluate a framework for assessing surgical video clips by categorizing them based on the surgical step being performed and the level of the surgeon's competence. Design, Setting, and Participants This quality improvement study assessed 103 video clips of 8 surgeons of various levels performing knot tying, suturing, and needle passing from the Johns Hopkins University-Intuitive Surgical Gesture and Skill Assessment Working Set. Data were collected before 2015, and data analysis took place from March to July 2019. Main Outcomes and Measures Deep learning models were trained to estimate categorical outputs such as performance level (ie, novice, intermediate, and expert) and surgical actions (ie, knot tying, suturing, and needle passing). The efficacy of these models was measured using precision, recall, and model accuracy. Results The provided architectures achieved accuracy in surgical action and performance calculation tasks using only video input. The embedding representation had a mean (root mean square error [RMSE]) precision of 1.00 (0) for suturing, 0.99 (0.01) for knot tying, and 0.91 (0.11) for needle passing, resulting in a mean (RMSE) precision of 0.97 (0.01). Its mean (RMSE) recall was 0.94 (0.08) for suturing, 1.00 (0) for knot tying, and 0.99 (0.01) for needle passing, resulting in a mean (RMSE) recall of 0.98 (0.01). It also estimated scores on the Objected Structured Assessment of Technical Skill Global Rating Scale categories, with a mean (RMSE) precision of 0.85 (0.09) for novice level, 0.67 (0.07) for intermediate level, and 0.79 (0.12) for expert level, resulting in a mean (RMSE) precision of 0.77 (0.04). Its mean (RMSE) recall was 0.85 (0.05) for novice level, 0.69 (0.14) for intermediate level, and 0.80 (0.13) for expert level, resulting in a mean (RMSE) recall of 0.78 (0.03). Conclusions and Relevance The proposed models and the accompanying results illustrate that deep machine learning can identify associations in surgical video clips. These are the first steps to creating a feedback mechanism for surgeons that would allow them to learn from their experiences and refine their skills.