The importance of being external. methodological insights for the external validation of machine learning models in medicine

The importance of being external. methodological insights for the external validation of machine learning models in medicine
复制标题

DOI:
10.1016/j.cmpb.2021.106288
复制
发表时间:
2021-08-02
影响因子:
6.1
通讯作者:
Carobene, Anna
Carobene, Anna
中科院分区:
工程技术2区
文献类型:
--
作者:
Cabitza, Federico;Campagner, Andrea;Carobene, Anna

文献摘要

被引文献

相似文献

背景与目的医学机器学习(ML)模型在同一队列的数据上往往比在新数据上表现得更好,这通常是由于过度拟合或协变量转移。因此,外部验证(EV)是医疗ML评估中的一种必要做法。然而,在如何解释EV结果从而评估ML模型的稳健性方面,文献中仍然存在空白。方法:我们提出了一种元验证方法来评估EV过程的稳健性。在这样做的过程中,我们通过考虑数据集的基数以及EV数据集与训练集的相似性来补充通常的评估EV的方法。然后,我们通过将基数和相似性的概念集成到两个总结性数据可视化中,研究如何使用基数和相似性的概念来告知验证过程的可靠性。结果:我们通过应用我们的方法对来自3个不同大陆的8辆电动汽车的最先进的新冠肺炎诊断模型进行验证,从而说明了我们的方法。模型性能受到数据相似性的适度影响(Pearson Rho=0.38,p<0.001)。在电动汽车中,经过验证的模型报告了良好的AUC(平均:0.84)、可接受的校准(平均:0.17)和实用性(平均:0.50)。验证数据集在数据集基数和相似性方面是足够的,因此表明结果是可靠的。结论:本文提出了一种新的精益方法:1)研究训练集和验证集之间的相似性如何影响最大似然模型的泛化能力;2)从判别、效用和校准三个互补的性能维度评估电动汽车评估的合理性;3)得出验证下模型的稳健性结论。我们将这一方法应用于从常规血液测试中诊断新冠肺炎的最新模型,并展示了如何根据所提出的框架来解释结果。(C)2021年爱思唯尔B.V.保留所有权利。
Background and Objective Medical machine learning (ML) models tend to perform better on data from the same cohort than on new data, often due to overfitting, or co-variate shifts. For these reasons, external validation (EV) is a necessary practice in the evaluation of medical ML. However, there is still a gap in the literature on how to interpret EV results and hence assess the robustness of ML models.Methods: We fill this gap by proposing a meta-validation method, to assess the soundness of EV procedures. In doing so, we complement the usual way to assess EV by considering both dataset cardinality, and the similarity of the EV dataset with respect to the training set. We then investigate how the notions of cardinality and similarity can be used to inform on the reliability of a validation procedure, by integrating them into two summative data visualizations.Results: We illustrate our methodology by applying it to the validation of a state-of-the-art COVID-19 diagnostic model on 8 EV sets, collected across 3 different continents. The model performance was moderately impacted by data similarity (Pearson rho = 0.38, p < 0.001). In the EV, the validated model reported good AUC (average: 0.84), acceptable calibration (average: 0.17) and utility (average: 0.50). The validation datasets were adequate in terms of dataset cardinality and similarity, thus suggesting the soundness of the results. We also provide a qualitative guideline to evaluate the reliability of validation procedures, and we discuss the importance of proper external validation in light of the obtained results.Conclusions: In this paper, we propose a novel, lean methodology to: 1) study how the similarity between training and validation sets impacts the generalizability of a ML model; 2) assess the soundness of EV evaluations along three complementary performance dimensions: discrimination, utility and calibration; 3) draw conclusions on the robustness of the model under validation. We applied this methodology to a state-of-the-art model for the diagnosis of COVID-19 from routine blood tests, and showed how to interpret the results in light of the presented framework. (C) 2021 Elsevier B.V. All rights reserved.