Improving the repeatability of deep learning models with Monte Carlo dropout.

Improving the repeatability of deep learning models with Monte Carlo dropout.
复制标题

DOI:
10.1038/s41746-022-00709-3
复制
发表时间:
2022-11-18
影响因子:
15.2
通讯作者:
--
中科院分区:
医学1区
文献类型:
--
作者:

文献摘要

参考文献

被引文献

相似文献

将人工智能集成到临床工作流程中需要可靠且稳健的模型。可重复性是模型稳健性的一个关键属性。理想的可重复模型在类似条件下进行的独立测试中输出预测没有变化。然而,轻微的变化虽然并不理想,但在实践中可能是不可避免的并且是可以接受的。在模型开发和评估过程中,过多关注分类性能,而很少评估模型的重复性,导致开发出的模型无法用于临床实践。在这项工作中,我们评估了四种模型类型(二元分类、多类分类、序数分类和回归)对同一患者在同一次就诊期间获取的图像的可重复性。我们研究了每个模型在来自公共和私人数据集的四种医学图像分类任务上的表现:膝骨关节炎、宫颈癌筛查、乳腺密度估计和早产儿视网膜病变。在 ResNet 和 DenseNet 架构上测量和比较可重复性。此外,我们评估了测试时采样蒙特卡洛丢失预测对分类性能和可重复性的影响。利用蒙特卡洛预测显着提高了二元、多类和序数模型上所有任务的可重复性,特别是在类边界处,导致 95% 一致限度平均降低 16%,类不一致率平均降低 7%。在大多数设置中,分类准确性以及可重复性都得到了提高。我们的结果表明,超过大约 20 次蒙特卡罗迭代后,重复性就没有进一步提高。除了更高的重测一致性之外,蒙特卡罗预测得到了更好的校准,这使得输出概率更准确地反映了正确分类的真实可能性。
The integration of artificial intelligence into clinical workflows requires reliable and robust models. Repeatability is a key attribute of model robustness. Ideal repeatable models output predictions without variation during independent tests carried out under similar conditions. However, slight variations, though not ideal, may be unavoidable and acceptable in practice. During model development and evaluation, much attention is given to classification performance while model repeatability is rarely assessed, leading to the development of models that are unusable in clinical practice. In this work, we evaluate the repeatability of four model types (binary classification, multi-class classification, ordinal classification, and regression) on images that were acquired from the same patient during the same visit. We study the each model’s performance on four medical image classification tasks from public and private datasets: knee osteoarthritis, cervical cancer screening, breast density estimation, and retinopathy of prematurity. Repeatability is measured and compared on ResNet and DenseNet architectures. Moreover, we assess the impact of sampling Monte Carlo dropout predictions at test time on classification performance and repeatability. Leveraging Monte Carlo predictions significantly increases repeatability, in particular at the class boundaries, for all tasks on the binary, multi-class, and ordinal models leading to an average reduction of the 95% limits of agreement by 16% points and of the class disagreement rate by 7% points. The classification accuracy improves in most settings along with the repeatability. Our results suggest that beyond about 20 Monte Carlo iterations, there is no further gain in repeatability. In addition to the higher test-retest agreement, Monte Carlo predictions are better calibrated which leads to output probabilities reflecting more accurately the true likelihood of being correctly classified.
DOI: 10.1002/mrm.28022
发表时间: 2019-10-21
影响因子: 3.3
作者:
Estrada, Santiago;Lu, Ran;Reuter, Martin
通讯作者: Reuter, Martin
DOI: 10.1148/ryai.2020190199
发表时间: 2021-01-01
期刊: RADIOLOGY-ARTIFICIAL INTELLIGENCE
影响因子: --
作者:
Hoebel, Katharina, V;Patel, Jay B.;Kalpathy-Cramer, Jayashree
通讯作者: Kalpathy-Cramer, Jayashree
DOI: 10.1093/annonc/mdy520
发表时间: 2019-02-01
期刊: Annals of oncology : official journal of the European Society for Medical Oncology
影响因子: --
作者:
Haenssle, H A;Fink, C;Uhlmann, L
通讯作者: Uhlmann, L
DOI: 10.1016/s2214-109x(19)30482-6
发表时间: 2020-02-01
影响因子: 34.3
作者:
Arbyn, Marc;Weiderpass, Elisabete;Bray, Freddie
通讯作者: Bray, Freddie
DOI: 10.1093/jnci/87.9.670
发表时间: 1995-05-03
期刊: JOURNAL OF THE NATIONAL CANCER INSTITUTE
影响因子: --
作者:
BOYD, NF;BYNG, JW;YAFFE, MJ
通讯作者: YAFFE, MJ