On the assessment of the added value of new predictive biomarkers.

On the assessment of the added value of new predictive biomarkers.
复制标题

DOI:
10.1186/1471-2288-13-98
复制
发表时间:
2013-07-29
影响因子:
4
通讯作者:
Petrick N
Petrick N
中科院分区:
医学3区
文献类型:
--
作者:
Chen W;Samuelson FW;Gallas BD;Kang L;Sahiner B;Petrick N

文献摘要

参考文献

被引文献

相似文献

生物标志物开发的激增要求对统计评估方法进行研究,以严格评估新兴的生物标志物和分类模型。最近,几位作者报告了一个令人困惑的观察结果,即在logistic回归模型中评估新生物标志物对现有生物标志物的附加值时,新预测变量的统计学显著性不一定转化为ROC曲线下面积(AUC)的统计学显著性增加。Vickers等人得出结论,这种不一致是因为AUC“具有非常差的统计特性”,即,这是非常保守的。该声明基于误用DeLong等人方法的模拟。我们的目的是提供一个公平的比较似然比(LR)测试和Wald测试与诊断准确性(AUC)测试。我们提出了一个测试来比较理想的AUC嵌套线性判别函数通过F检验。并与Logistic回归模型的LR检验和Wald检验进行了比较。这三个检验的零假设是等效的;但是,F检验是精确检验,而LR检验和Wald检验是渐近检验。我们的模拟表明,F检验具有标称的I类错误,即使在一个小的样本量。我们的研究结果还表明,LR测试和Wald测试膨胀的I型错误时,样本容量是小的,而I型错误渐近收敛到标称值随着样本容量的增加,正如预期的那样。我们进一步表明,DeLong等人。方法测试一个不同的假设,并有名义上的I型错误时,它是在其设计范围内使用。最后,我们总结了本文考虑的所有四种方法的优缺点。我们表明,对于显示新生物标志物的有用性或表征分类模型的性能,ROC分析没有什么内在的不那么强大或不一致。用于评估生物标志物和分类模型的每种统计方法都有其自身的优点和缺点。研究者需要根据评估目的、正在进行评估的生物标志物开发阶段、可用的患者数据以及方法学背后假设的有效性来选择方法。
The surge in biomarker development calls for research on statistical evaluation methodology to rigorously assess emerging biomarkers and classification models. Recently, several authors reported the puzzling observation that, in assessing the added value of new biomarkers to existing ones in a logistic regression model, statistical significance of new predictor variables does not necessarily translate into a statistically significant increase in the area under the ROC curve (AUC). Vickers et al. concluded that this inconsistency is because AUC “has vastly inferior statistical properties,” i.e., it is extremely conservative. This statement is based on simulations that misuse the DeLong et al. method. Our purpose is to provide a fair comparison of the likelihood ratio (LR) test and the Wald test versus diagnostic accuracy (AUC) tests. We present a test to compare ideal AUCs of nested linear discriminant functions via an F test. We compare it with the LR test and the Wald test for the logistic regression model. The null hypotheses of these three tests are equivalent; however, the F test is an exact test whereas the LR test and the Wald test are asymptotic tests. Our simulation shows that the F test has the nominal type I error even with a small sample size. Our results also indicate that the LR test and the Wald test have inflated type I errors when the sample size is small, while the type I error converges to the nominal value asymptotically with increasing sample size as expected. We further show that the DeLong et al. method tests a different hypothesis and has the nominal type I error when it is used within its designed scope. Finally, we summarize the pros and cons of all four methods we consider in this paper. We show that there is nothing inherently less powerful or disagreeable about ROC analysis for showing the usefulness of new biomarkers or characterizing the performance of classification models. Each statistical method for assessing biomarkers and classification models has its own strengths and weaknesses. Investigators need to choose methods based on the assessment purpose, the biomarker development phase at which the assessment is being performed, the available patient data, and the validity of assumptions behind the methodologies.
关键评估用于分类或预测的生物标志物的准确性:研究设计标准。
DOI: 10.1093/jnci/djn326
发表时间: 2008-10-15
影响因子: 10.3
作者:
Pepe, Margaret S.;Feng, Ziding;Janes, Holly;Bossuyt, Patrick M.;Potter, John D.
通讯作者: Potter, John D.
DOI: 10.1002/sim.5727
发表时间: 2013-04-30
影响因子: 2
作者:
Pepe, Margaret Sullivan;Kerr, Kathleen F.;Longton, Gary;Wang, Zheyu
通讯作者: Wang, Zheyu
DOI: 10.1214/aoms/1177730196
发表时间: 1948-01-01
影响因子: --
作者:
HOEFFDING, W
通讯作者: HOEFFDING, W
DOI: 10.1118/1.2868757
发表时间: 2008-04-01
期刊: MEDICAL PHYSICS
影响因子: 3.8
作者:
Sahiner, Berkman;Chan, Heang-Ping;Hadjiiski, Lubomir
通讯作者: Hadjiiski, Lubomir
DOI: 10.2307/2332629
发表时间: 1948-01-01
期刊: BIOMETRIKA
影响因子: 2.7
作者:
RAO, CR
通讯作者: RAO, CR