Generalizability of a Machine Learning Model for Improving Utilization of Parathyroid Hormone-Related Peptide Testing across Multiple Clinical Centers.

Generalizability of a Machine Learning Model for Improving Utilization of Parathyroid Hormone-Related Peptide Testing across Multiple Clinical Centers.
复制标题

机器学习模型的通用性,可提高多个临床中心甲状旁腺激素相关肽测试的利用率。

DOI:
10.1093/clinchem/hvad141
复制
发表时间:
2023
期刊:
影响因子:
9.3
通讯作者:
Wang,Fei
Wang,Fei
中科院分区:
医学1区
文献类型:
--
作者:
Yang,HeS;Pan,Weishen;Wang,Yingheng;Zaydman,MarkA;Spies,NicholasC;Zhao,Zhen;Guise,TheresaA;Meng,QingH;Wang,Fei

文献摘要

被引文献

相似文献

背景测量甲状旁腺激素相关肽(PTHrP)有助于诊断恶性肿瘤的体液高钙血症,但通常是针对预检测可能性较低的患者进行的,导致测试利用率较差。手动审查结果,以确定不适当的PTHrP orders是一个繁琐的process.MethodsUsing从一个单一的institute的1330例患者的数据集,我们开发了一个机器学习(ML)模型来预测异常PTHrP的结果。然后,我们在两个外部数据集上评估了模型的性能。不同的策略(模型传输,再训练,重建和微调)进行了研究,以提高模型的泛化能力。采用最大平均差异(MMD)来量化不同dataset.ResultsThe模型的数据分布的转变,实现了0.936的受试者工作特征曲线(AUROC)下的面积,在0.900灵敏度在发展队列的特异性为0.842。直接将该模型传输到两个外部数据集导致AUROC恶化至0.838和0.737,后者具有更大的MMD,与原始数据集相比,对应于更大的数据偏移。模型重建使用特定网站的数据提高AUROC为0.891和0.837的两个网站,分别。当外部数据不足以进行再培训,微调策略也提高了模型utility.ConclusionsML提供承诺,以提高PTHrP测试利用率,同时减轻人工审查的负担。将现成的模型传输到外部数据集可能会由于数据分布偏移而导致性能恶化。当有足够的数据时,模型再训练或重建可以提高泛化能力,当特定于站点的数据有限时,模型微调可能是有利的。
BackgroundMeasuring parathyroid hormone-related peptide (PTHrP) helps diagnose the humoral hypercalcemia of malignancy, but is often ordered for patients with low pretest probability, resulting in poor test utilization. Manual review of results to identify inappropriate PTHrP orders is a cumbersome process.MethodsUsing a dataset of 1330 patients from a single institute, we developed a machine learning (ML) model to predict abnormal PTHrP results. We then evaluated the performance of the model on two external datasets. Different strategies (model transporting, retraining, rebuilding, and fine-tuning) were investigated to improve model generalizability. Maximum mean discrepancy (MMD) was adopted to quantify the shift of data distributions across different datasets.ResultsThe model achieved an area under the receiver operating characteristic curve (AUROC) of 0.936, and a specificity of 0.842 at 0.900 sensitivity in the development cohort. Directly transporting this model to two external datasets resulted in a deterioration of AUROC to 0.838 and 0.737, with the latter having a larger MMD corresponding to a greater data shift compared to the original dataset. Model rebuilding using site-specific data improved AUROC to 0.891 and 0.837 on the two sites, respectively. When external data is insufficient for retraining, a fine-tuning strategy also improved model utility.ConclusionsML offers promise to improve PTHrP test utilization while relieving the burden of manual review. Transporting a ready-made model to external datasets may lead to performance deterioration due to data distribution shift. Model retraining or rebuilding could improve generalizability when there are enough data, and model fine-tuning may be favorable when site-specific data is limited.