Disadvantages of using the area under the receiver operating characteristic curve to assess imaging tests: A discussion and proposal for an alternative approach

Disadvantages of using the area under the receiver operating characteristic curve to assess imaging tests: A discussion and proposal for an alternative approach
复制标题

DOI:
10.1007/s00330-014-3487-0
复制
发表时间:
2015-04-01
期刊:
影响因子:
5.9
通讯作者:
Mallett, Susan
Mallett, Susan
中科院分区:
医学2区
文献类型:
--
作者:
Halligan, Steve;Altman, Douglas G.;Mallett, Susan

文献摘要

被引文献

相似文献

目的是描述受试者工作特征曲线下面积(ROC AUC)在衡量诊断测试性能方面的不足,并提出基于净收益的替代方案。我们使用叙述性综述,并补充了一项计算机辅助检测CT结肠成像研究的数据。我们发现ROC AUC存在问题。读者信心得分呈高度非正态分布,得分分布呈双峰分布。因此,ROC曲线被高度外推,AUC主要依赖于没有患者数据的区域。AUC取决于所用的曲线拟合方法。ROC AUC不考虑由假阴性和假阳性诊断引起的流行率或不同的误分类成本。ROC AUC的变化对临床医生的直接临床意义不大。基于临床相关阈值下敏感度和特异度的变化,提出了基于净收益的替代分析。净收益结合了对患病率和错误分类成本的估计,它在临床上是可解释的,因为它反映了当引入新的诊断测试时正确和错误诊断的变化。ROC AUC在测试评估的早期阶段最有用,而基于净收益的方法在已知临床背景的情况下评估放射学测试更有用。净效益更有助于评估临床效果。受试者工作特征曲线下面积(ROC AUC)衡量诊断的准确性。用于建立ROC曲线的可信分数可能难以分配。假阳性和假阴性诊断有不同的误判代价。过度的ROC曲线外推是不可取的。净效益方法可能比ROC AUC提供更有意义和临床可解释的结果。
The objectives are to describe the disadvantages of the area under the receiver operating characteristic curve (ROC AUC) to measure diagnostic test performance and to propose an alternative based on net benefit.We use a narrative review supplemented by data from a study of computer-assisted detection for CT colonography.We identified problems with ROC AUC. Confidence scoring by readers was highly non-normal, and score distribution was bimodal. Consequently, ROC curves were highly extrapolated with AUC mostly dependent on areas without patient data. AUC depended on the method used for curve fitting. ROC AUC does not account for prevalence or different misclassification costs arising from false-negative and false-positive diagnoses. Change in ROC AUC has little direct clinical meaning for clinicians. An alternative analysis based on net benefit is proposed, based on the change in sensitivity and specificity at clinically relevant thresholds. Net benefit incorporates estimates of prevalence and misclassification costs, and it is clinically interpretable since it reflects changes in correct and incorrect diagnoses when a new diagnostic test is introduced.ROC AUC is most useful in the early stages of test assessment whereas methods based on net benefit are more useful to assess radiological tests where the clinical context is known. Net benefit is more useful for assessing clinical impact.The area under the receiver operating characteristic curve (ROC AUC) measures diagnostic accuracy.Confidence scores used to build ROC curves may be difficult to assign.False-positive and false-negative diagnoses have different misclassification costs.Excessive ROC curve extrapolation is undesirable.Net benefit methods may provide more meaningful and clinically interpretable results than ROC AUC.