AI in the Loop: functionalizing fold performance disagreement to monitor automated medical image segmentation workflows.

AI in the Loop: functionalizing fold performance disagreement to monitor automated medical image segmentation workflows.
复制标题

DOI:
10.3389/fradi.2023.1223294
复制
发表时间:
2023
期刊:
Frontiers in radiology
影响因子:
--
通讯作者:
Kline, Timothy L.
Kline, Timothy L.
中科院分区:
其他
文献类型:
--
作者:
Gottlich, Harrison C.;Korfiatis, Panagiotis;Gregory, Adriana V.;Kline, Timothy L.

文献摘要

参考文献

相似文献

迫切需要自动标记表现不佳的预测的方法,以安全地将机器学习工作流程应用于临床实践,并在模型训练期间识别困难病例。使用折叠之间的骰子得分量化五重交叉验证子模型之间的不一致,并总结为模型置信度的替代。将汇总的Interfold Dices与人类观察者间值通知的阈值进行比较,以确定是否应手动审查最终集成模型性能。该方法在所有任务上都有效地标记了不好的分割图像,而无需参考标准。使用中值折叠间骰子进行比较,在排除标记图像后,观察到域内CT(0.85 ± 0.20至0.91 ± 0.08,8/50图像标记)和MR(0.76 ± 0.27至0.85 ± 0.09,8/50图像标记)的骰子评分显著改善。最令人印象深刻的是,在模拟的分布外任务中,模型在根治性肾切除术数据集上进行训练,不同的造影剂相位预测部分肾切除术所有皮质-髓质相位数据集(0.67 ± 0.36至0.89 ± 0.10,122/300图像标记)。当参考标准不可用时,将交叉子模型不一致与人类观察者间值进行比较是评估自动预测的有效且高效的方法。此功能为患者护理提供了必要的保障,对于安全实施自动化医学图像分割工作流程至关重要。
Methods that automatically flag poor performing predictions are drastically needed to safely implement machine learning workflows into clinical practice as well as to identify difficult cases during model training. Disagreement between the fivefold cross-validation sub-models was quantified using dice scores between folds and summarized as a surrogate for model confidence. The summarized Interfold Dices were compared with thresholds informed by human interobserver values to determine whether final ensemble model performance should be manually reviewed. The method on all tasks efficiently flagged poor segmented images without consulting a reference standard. Using the median Interfold Dice for comparison, substantial dice score improvements after excluding flagged images was noted for the in-domain CT (0.85 ± 0.20 to 0.91 ± 0.08, 8/50 images flagged) and MR (0.76 ± 0.27 to 0.85 ± 0.09, 8/50 images flagged). Most impressively, there were dramatic dice score improvements in the simulated out-of-distribution task where the model was trained on a radical nephrectomy dataset with different contrast phases predicting a partial nephrectomy all cortico-medullary phase dataset (0.67 ± 0.36 to 0.89 ± 0.10, 122/300 images flagged). Comparing interfold sub-model disagreement against human interobserver values is an effective and efficient way to assess automated predictions when a reference standard is not available. This functionality provides a necessary safeguard to patient care important to safely implement automated medical image segmentation workflows.
DOI: 10.1371/journal.pmed.1002689
发表时间: 2018-11
期刊: PLoS medicine
影响因子: 15.8
作者:
Vayena E;Blasimme A;Cohen IG
通讯作者: Cohen IG
DOI: 10.1016/j.media.2022.102680
发表时间: 2023-02
影响因子: 10.9
作者:
Bilic, Patrick;Christ, Patrick;Li, Hongwei Bran;Vorontsov, Eugene;Ben-Cohen, Avi;Kaissis, Georgios;Szeskin, Adi;Jacobs, Colin;Mamani, Gabriel Efrain Humpire;Chartrand, Gabriel;Lohoefer, Fabian;Holch, Julian Walter;Sommer, Wieland;Hofmann, Felix;Hostettler, Alexandre;Lev-Cohain, Naama;Drozdzal, Michal;Amitai, Michal Marianne;Vivanti, Refael;Sosna, Jacob;Ezhov, Ivan;Sekuboyina, Anjany;Navarro, Fernando;Kofler, Florian;Paetzold, Johannes C.;Shit, Suprosanna;Hu, Xiaobin;Lipkova, Jana;Rempfler, Markus;Piraud, Marie;Kirschke, Jan;Wiestler, Benedikt;Zhang, Zhiheng;Huelsemeyer, Christian;Beetz, Marcel;Ettlinger, Florian;Antonelli, Michela;Bae, Woong;Bellver, Miriam;Bi, Lei;Chen, Hao;Chlebus, Grzegorz;Dam, Erik B.;Dou, Qi;Fu, Chi-Wing;Georgescu, Bogdan;Giro-I-Nieto, Xavier;Gruen, Felix;Han, Xu;Heng, Pheng-Ann;Hesser, Jurgen;Moltz, Jan Hendrik;Igel, Christian;Isensee, Fabian;Jaeger, Paul;Jia, Fucang;Kaluva, Krishna Chaitanya;Khened, Mahendra;Kim, Ildoo;Kim, Jae-Hun;Kim, Sungwoong;Kohl, Simon;Konopczynski, Tomasz;Kori, Avinash;Krishnamurthi, Ganapathy;Li, Fan;Li, Hongchao;Li, Junbo;Li, Xiaomeng;Lowengrub, John;Ma, Jun;Maier-Hein, Klaus;Maninis, Kevis-Kokitsi;Meine, Hans;Merhof, Dorit;Pai, Akshay;Perslev, Mathias;Petersen, Jens;Pont-Tuset, Jordi;Qi, Jin;Qi, Xiaojuan;Rippel, Oliver;Roth, Karsten;Sarasua, Ignacio;Schenk, Andrea;Shen, Zengming;Torres, Jordi;Wachinger, Christian;Wang, Chunliang;Weninger, Leon;Wu, Jianrong;Xu, Daguang;Yang, Xiaoping;Yu, Simon Chun-Ho;Yuan, Yading;Yue, Miao;Zhang, Liping;Cardoso, Jorge;Bakas, Spyridon;Braren, Rickmer;Heinemann, Volker;Pal, Christopher;Tang, An;Kadoury, Samuel;Soler, Luc;van Ginneken, Bram;Greenspan, Hayit;Joskowicz, Leo;Menze, Bjoern
通讯作者: Menze, Bjoern
DOI: 10.1681/asn.2020040449
发表时间: 2020-11-01
影响因子: 13.6
作者:
Denic, Aleksandar;Elsherbiny, Hisham;Rule, Andrew D.
通讯作者: Rule, Andrew D.
DOI: 10.2196/13659
发表时间: 2019-07-10
影响因子: 7.4
作者:
Shaw, James;Rudzicz, Frank;Goldfarb, Avi
通讯作者: Goldfarb, Avi
DOI: 10.1117/1.jmi.6.3.034001
发表时间: 2019-07-01
影响因子: 2.4
作者:
Mueller, Sabine;Farag, Iva;Graf, Norbert
通讯作者: Graf, Norbert