Road to automating robotic suturing skills assessment: Battling mislabeling of the ground truth.

Road to automating robotic suturing skills assessment: Battling mislabeling of the ground truth.
复制标题

DOI:
10.1016/j.surg.2021.08.014
复制
发表时间:
2022-04
期刊:
影响因子:
3.8
通讯作者:
Liu Y
Liu Y
中科院分区:
医学2区
文献类型:
--
作者:
Hung AJ;Rambhatla S;Sanford DI;Pachauri N;Vanstrum E;Nguyen JH;Liu Y

文献摘要

参考文献

被引文献

相似文献

使用机器人器械运动学数据自动化外科医生技能评估。此外,实现一种无监督的错误标记检测算法,以识别可能被错误标记的样本,这些样本可以被删除以提高模型性能。视频记录和器械运动学数据来自于在Mimic FlexVR™机器人模拟器上完成的缝合练习。开发了一个结构化的人类共识建立过程,以确定机器人吻合能力评估(RACE)技术分数跨越三个人类年级。基于两层lstm(长短期记忆)分类模型,利用器械运动数据自动评估缝合技能。使用无监督标签分析仪(NoiseRank)来识别潜在的错误标记技能数据。采用最佳曲线下面积(AUC)来衡量LSTM模型的技能分数预测效果。NoiseRank根据错误标注的可能性输出了一份评级技能评估的排名列表。22名外科医生进行了226次缝合尝试,分为1404个个人技能评估点。利用所有可用数据,自动进针角度、进针和退针技术技能得分(AUC 0.698 - 0.705)优于基线定位(0.532)。随后,NoiseRank识别并删除了潜在的错误标签,提高了所有领域的模型性能(AUC为0.551 - 0.766)。机器学习模型使用来自人类评分者和机器人仪器运动学数据的地面真实值标签,可以自动评估详细的缝合技术技能,并具有良好的性能。此外,一种无监督的错误标记检测算法预测了错误标记的数据,允许它们的去除和随后的模型性能改进。
To automate surgeon skills evaluation using robotic instrument kinematic data. Additionally, to implement an unsupervised mislabeling detection algorithm to identify potentially mislabeled samples that can be removed to improve model performance. Video recordings and instrument kinematic data were derived from suturing exercises completed on the Mimic FlexVR™ robotic simulator. A structured human consensus-building process was developed to determine Robotic Anastomosis Competency Evaluation (RACE) technical scores across three human graders. A two-layer LSTM-based (long short-term memory) classification model used instrument kinematic data to automate suturing skills assessment. An unsupervised label analyzer (NoiseRank) was used to identify potential mislabeling of skills data. Performance of the LSTM model’s technical skill score prediction was measured by best area under the curve (AUC) over the training runs. NoiseRank outputted a ranked list of rated skills assessments based on likelihood of mislabeling. 22 surgeons performed 226 suturing attempts, which were broken down into 1,404 individual skill assessment points. Automation of Needle Entry Angle, Needle Driving, and Needle Withdrawal technical skill scores performed better (AUC 0.698 – 0.705) than Needle Positioning (0.532) at baseline utilizing all available data. Potential mislabels were subsequently identified by NoiseRank and removed, improving model performance across all domains (AUC 0.551 – 0.766). Using ground truth labels from human graders and robotic instrument kinematic data, machine learning models have automated assessment of detailed suturing technical skills with good performance. Further, an unsupervised mislabeling detection algorithm projected mislabeled data, allowing for their removal and subsequent improvement of model performance.
DOI: 10.1136/bmjopen-2014-006759
发表时间: 2015-06-15
期刊: BMJ open
影响因子: 2.9
作者:
Trehan A;Barnett-Vanes A;Carty MJ;McCulloch P;Maruthappu M
通讯作者: Maruthappu M
DOI: 10.1111/bju.14599
发表时间: 2019-05-01
期刊: BJU INTERNATIONAL
影响因子: 4.5
作者:
Hung, Andrew J.;Oh, Paul J.;Gill, Inderbir S.
通讯作者: Gill, Inderbir S.
DOI: 10.1016/j.euf.2021.04.001
发表时间: 2022-03
影响因子: 5.4
作者:
Trinh L;Mingo S;Vanstrum EB;Sanford DI;Aastha;Ma R;Nguyen JH;Liu Y;Hung AJ
通讯作者: Hung AJ
DOI: 10.1016/j.juro.2011.09.032
发表时间: 2012-01-01
期刊: JOURNAL OF UROLOGY
影响因子: 6.6
作者:
Goh, Alvin C.;Goldfarb, David W.;Dunkin, Brian J.
通讯作者: Dunkin, Brian J.
DOI: 10.1002/bjs.1800840237
发表时间: 1997-02-01
影响因子: 9.6
作者:
Martin, JA;Regehr, G;Brown, M
通讯作者: Brown, M