Bias in sensitivity and specificity caused by data-driven selection of optimal cutoff values: Mechanisms, magnitude, and solutions

Bias in sensitivity and specificity caused by data-driven selection of optimal cutoff values: Mechanisms, magnitude, and solutions
复制标题

DOI:
10.1373/clinchem.2007.096032
复制
发表时间:
2008-04-01
期刊:
影响因子:
9.3
通讯作者:
Zwinderman, Aielko H.
Zwinderman, Aielko H.
中科院分区:
医学1区
文献类型:
--
作者:
Leeflang, Mariska M. G.;Moons, Karel G. M.;Zwinderman, Aielko H.

文献摘要

被引文献

相似文献

背景:涉及连续变量的测试结果的最佳截止值通常是以数据驱动的方式得出的。然而,这种方法可能导致诊断准确性的过度乐观的措施。我们评估了与数据驱动的选择截止值的敏感性和特异性的偏差的大小,并研究了潜在的解决方案,以减少这种bias.METHODS:不同的样本量,分布和患病率被用于模拟研究。我们比较了基于Youden指数的数据驱动的准确性估计值与真实值,并计算了中位数偏差。三种替代方法(假设一个特定的分布,留一法,平滑的ROC曲线)进行了检查,以减少这种bias.RESULTS的能力:由数据驱动的截止值的优化所造成的偏差的大小是负相关的样本量。如果灵敏度和特异性的真实值均为84%,则样本量为40的研究中的估计值约为90%。如果样本量增加到200,估计值将为86%。当样本量保持不变时,测试结果的分布对偏倚量的影响很小。更强大的方法,优化截止值不太容易出现偏差,但性能恶化,如果基本的假设是不满足.CONCLUSIONS:数据驱动的最佳截止值的选择可能会导致过于乐观的估计灵敏度和特异性,特别是在小型研究。替代方法可以减少这种偏差,但找到截断值和准确性的可靠估计需要相当大的样本量。(c)2008年美国临床化学协会。
BACKGROUND: Optimal cutoff values for tests results involving continuous variables are often derived in a data-driven way. This approach, however, may lead to overly optimistic measures of diagnostic accuracy. We evaluated the magnitude of the bias in sensitivity and specificity associated with data-driven selection of cutoff values and examined potential solutions to reduce this bias.METHODS: Different sample sizes, distributions, and prevalences were used in a simulation study. We compared data-driven estimates of accuracy based on the Youden index with the true values and calculated the median bias. Three alternative approaches (assuming a specific distribution, leave-one-out, smoothed ROC curve) were examined for their ability to reduce this bias.RESULTS: The magnitude of bias caused by data-driven optimization of cutoff values was inversely related to sample size. If the true values for sensitivity and specificity are both 84%, the estimates in studies with a sample size of 40 will be approximately 90%. If the sample size increases to 200, the estimates will be 86%. The distribution of the test results had little impact on the amount of bias when sample size was held constant. More robust methods of optimizing cutoff values were less prone to bias, but the performance deteriorated if the underlying assumptions were not met.CONCLUSIONS: Data-driven selection of the optimal cutoff value can lead to overly optimistic estimates of sensitivity and specificity, especially in small studies. Alternative methods can reduce this bias, but finding robust estimates for cutoff values and accuracy requires considerable sample sizes. (c) 2008 American Association for Clinical Chemistry.