Commentary: Reference-test bias in diagnostic-test evaluation: a problem for epidemiologists, too.

Commentary: Reference-test bias in diagnostic-test evaluation: a problem for epidemiologists, too.
复制标题

评论:诊断测试评估中的参考测试偏差:对于流行病学家来说也是一个问题。

DOI:
10.1097/ede.0b013e31823b5b5b
复制
发表时间:
2012
期刊:
Epidemiology (Cambridge, Mass.)
影响因子:
--
通讯作者:
Miller,WilliamC
Miller,WilliamC
中科院分区:
--
文献类型:
--
作者:
Miller,WilliamC

文献摘要

被引文献

相似文献

流行病学家在其工作的几乎所有方面都依赖于对疾病状态的准确评估,无论是研究还是实践。大多数流行病学家都知道诊断测试是容易出错的。他们开发并应用了复杂的方法来解决结果的测量误差。流行病学家对敏感性和特异性(用于表达诊断和筛查测试准确性的参数)感到满意。然而,尽管诊断测试在他们的工作中占据核心地位,但很少有“传统流行病学家”对评估诊断测试的方法做出了贡献。相反,诊断测试评估一直是“临床流行病学家”和少数生物统计学家的职权范围。诊断测试评估存在许多潜在的偏差。 1 从根本上讲,对新诊断测试的评估需要与参考(“金”)标准进行比较,通常认为该标准可以完美地区分疾病和非疾病状态。不幸的是,很少(如果有的话)参考标准是完美的。由此产生的参考测试偏差是最重要、最普遍、最具挑战性的偏差形式之一。 2, 3 参考测试偏差的影响可以用一个简单的例子来描述。考虑一个灵敏度为 0.85、特异性为 0.90 的参考测试,以及一个真实(但未知)灵敏度为 0.90、特异性为 0.95 的新的改进测试。为简单起见,我们假设这些测试的条件独立性。在患病率为 0.1 的研究样本中,新测试相对于参考测试的测量灵敏度和特异性将分别为 0.46 和 0.93。在患病率为 0.5 的研究样本中(即独立选择的病例和非病例),新测试的敏感性和特异性估计值分别显着提高至 0.81 和 0.83。敏感性和特异性估计值对患病率的依赖性与阳性和阴性预测值随患病率的变化相当。随着流行率的降低,参考测试的假阳性率也会增加;随着患病率的增加,参考测试假阴性也会增加。参考标准对疾病的不完美分类导致新测试明显错误分类,并导致敏感性和特异性估计出现偏差。当新测试优于参考标准时,参考测试偏差尤其成问题,如上述示例所示。人们相信新的测试本质上更好,这导致了直观的“解决方案”来解释参考测试偏差。不幸的是,与流行病学的其他领域一样,直觉往往是一个糟糕的统计学家。 20 世纪 90 年代,微生物学家采用了一种直观吸引人的程序,称为差异分析,用于评估用于诊断衣原体感染和其他传染病的新核酸扩增测试(例如聚合酶链反应)。 4-6 微生物学家认识到,这些新测试代表了相对于培养技术的重大进步,而已知培养技术的灵敏度有限。 7 为了解决使用培养物作为参考标准时参考测试偏差的问题,微生物
Epidemiologists depend on accurate assessment of disease states for almost all aspects of their work, whether for research or practice. Most epidemiologists are aware that diagnostic tests are fallible. They have developed and applied sophisticated methods to address the measurement error of outcomes. Epidemiologists are comfortable with sensitivity and specificity, the parameters used to express the accuracy of diagnostic and screening tests. However, few “traditional epidemiologists” have contributed to methods for evaluating diagnostic tests, despite the centrality of diagnostic tests to their work. Instead, diagnostic test evaluation has been the purview of “clinical epidemiologists” and a few biostatisticians.Diagnostic-test evaluation is subject to numerous potential biases. 1 Fundamentally, the evaluation of a new diagnostic test requires a comparison with a reference (“gold”) standard, which is usually assumed to discriminate disease and nondisease states perfectly. Unfortunately, few (if any) reference standards are perfect. The resulting reference-test bias is one of the most important, pervasive, and challenging forms of bias. 2, 3 The impact of reference-test bias can be described with a simple example. Consider a reference test with sensitivity of 0.85 and specificity of 0.90 and a new and improved test with a true (but unknown) sensitivity of 0.90 and specificity of 0.95. For simplicity, we assume conditional independence of these tests. In a study sample with a prevalence of 0.1, the measured sensitivity and specificity of the new test against the reference test would be 0.46 and 0.93, respectively. In a study sample with a prevalence of 0.5 (ie, cases and noncases selected independently), the estimates of sensitivity and specificity for the new test improve markedly to 0.81 and 0.83, respectively. The dependence of the sensitivity and specificity estimates on prevalence is comparable with the variation of positive and negative predictive values with prevalence. False positives by the reference test increase as prevalence decreases; reference-test false negatives increase as prevalence increases. The imperfect classification of disease by the reference standard leads to apparent misclassification by the new test and to biased sensitivity and specificity estimates. Reference-test bias is particularly problematic when the new test is better than the reference standard, as in the aforementioned example. The belief that a new test is inherently better has led to intuitive “solutions” to account for reference-test bias. Unfortunately, as with other areas of epidemiology, intuition is often a poor statistician. In the 1990s, microbiologists adopted an intuitively appealing procedure referred to as discrepant analysis for the evaluation of new nucleic-acid-amplification tests (such as polymerase chain reaction) for the diagnosis of chlamydial infection and other infectious diseases. 4–6 Microbiologists recognized that these new tests represented a major advance over culture techniques, which were known to have limited sensitivity. 7 To address the concern of reference-test bias when using culture as the reference standard, the microbi-