Empirical assessment of bias in machine learning diagnostic test accuracy studies

Empirical assessment of bias in machine learning diagnostic test accuracy studies
复制标题

DOI:
10.1093/jamia/ocaa075
复制
发表时间:
2020-07-01
影响因子:
6.4
通讯作者:
Ioannidis, John P. A.
Ioannidis, John P. A.
中科院分区:
管理学2区
文献类型:
--
作者:
Crowley, Ryan J.;Tan, Yuan Jin;Ioannidis, John P. A.

文献摘要

被引文献

相似文献

机器学习(ML)诊断工具在改善医疗保健方面具有巨大潜力。然而,方法上的缺陷可能会影响用于评估这些工具的诊断测试准确性研究。我们的目的是评价文献中设计特征的流行率和报告。此外,我们试图凭经验评估设计特点是否可能与不同的估计诊断accuracy.Materials and Methods:我们系统地检索2 × 2表(n=281)描述的性能ML诊断工具,来自114出版物在38荟萃分析,从PubMed。提取的数据包括测试性能、样本量和设计特征。一个混合效应的荟萃回归进行量化设计功能和诊断准确性之间的关联。结果:参与者的种族和测试解释盲未报告的90%和60%的研究,分别。报告偶尔缺乏基本特征,如研究设计(28%未报告)。44%的研究使用了没有适当保障措施的内部验证。几个设计特征与更大的准确性估计相关,包括未报告的(相对诊断比值比[RDOR],2.11; 95%置信区间[CI],1.43-3.1)或病例对照研究设计(RDOR,1.27; 95%CI,0.97-1.66),并招募参与者进行指数测试(RDOR,1.67; 95%CI,1.08-2.59)。研究设计特点可能会影响估计的ML诊断测试准确性literature.Conclusions:本研究确定的陷阱,威胁ML诊断工具的有效性,普遍性和临床价值,并提供改进建议。
Objective: Machine learning (ML) diagnostic tools have significant potential to improve health care. However, methodological pitfalls may affect diagnostic test accuracy studies used to appraise such tools. We aimed to evaluate the prevalence and reporting of design characteristics within the literature. Further, we sought to empirically assess whether design features may be associated with different estimates of diagnostic accuracy.Materials and Methods: We systematically retrieved 2 x 2 tables (n=281) describing the performance of ML diagnostic tools, derived from 114 publications in 38 meta-analyses, from PubMed. Data extracted included test performance, sample sizes, and design features. A mixed-effects metaregression was run to quantify the association between design features and diagnostic accuracy.Results: Participant ethnicity and blinding in test interpretation was unreported in 90% and 60% of studies, respectively. Reporting was occasionally lacking for rudimentary characteristics such as study design (28% unreported). Internal validation without appropriate safeguards was used in 44% of studies. Several design features were associated with larger estimates of accuracy, including having unreported (relative diagnostic odds ratio [RDOR], 2.11; 95% confidence interval [CI], 1.43-3.1) or case-control study designs (RDOR, 1.27; 95% CI, 0.97-1.66), and recruiting participants for the index test (RDOR, 1.67; 95% CI, 1.08-2.59).Discussion: Significant underreporting of experimental details was present. Study design features may affect estimates of diagnostic performance in the ML diagnostic test accuracy literature.Conclusions: The present study identifies pitfalls that threaten the validity, generalizability, and clinical value of ML diagnostic tools and provides recommendations for improvement.