Commentary: Reference-test bias in diagnostic-test evaluation: a problem for epidemiologists, too.
Commentary: Reference-test bias in diagnostic-test evaluation: a problem for epidemiologists, too.
复制标题
评论:诊断测试评估中的参考测试偏差:对于流行病学家来说也是一个问题。
DOI:
10.1097/ede.0b013e31823b5b5b
复制
发表时间:
2012
期刊:
影响因子:
--
通讯作者:
Miller,WilliamC
中科院分区:
文献类型:
--
作者:
Miller,WilliamC
Epidemiologists depend on accurate assessment of disease states for almost all aspects of their work, whether for research or practice. Most epidemiologists are aware that diagnostic tests are fallible. They have developed and applied sophisticated methods to address the measurement error of outcomes. Epidemiologists are comfortable with sensitivity and specificity, the parameters used to express the accuracy of diagnostic and screening tests. However, few “traditional epidemiologists” have contributed to methods for evaluating diagnostic tests, despite the centrality of diagnostic tests to their work. Instead, diagnostic test evaluation has been the purview of “clinical epidemiologists” and a few biostatisticians.Diagnostic-test evaluation is subject to numerous potential biases. 1 Fundamentally, the evaluation of a new diagnostic test requires a comparison with a reference (“gold”) standard, which is usually assumed to discriminate disease and nondisease states perfectly. Unfortunately, few (if any) reference standards are perfect. The resulting reference-test bias is one of the most important, pervasive, and challenging forms of bias. 2, 3 The impact of reference-test bias can be described with a simple example. Consider a reference test with sensitivity of 0.85 and specificity of 0.90 and a new and improved test with a true (but unknown) sensitivity of 0.90 and specificity of 0.95. For simplicity, we assume conditional independence of these tests. In a study sample with a prevalence of 0.1, the measured sensitivity and specificity of the new test against the reference test would be 0.46 and 0.93, respectively. In a study sample with a prevalence of 0.5 (ie, cases and noncases selected independently), the estimates of sensitivity and specificity for the new test improve markedly to 0.81 and 0.83, respectively. The dependence of the sensitivity and specificity estimates on prevalence is comparable with the variation of positive and negative predictive values with prevalence. False positives by the reference test increase as prevalence decreases; reference-test false negatives increase as prevalence increases. The imperfect classification of disease by the reference standard leads to apparent misclassification by the new test and to biased sensitivity and specificity estimates. Reference-test bias is particularly problematic when the new test is better than the reference standard, as in the aforementioned example. The belief that a new test is inherently better has led to intuitive “solutions” to account for reference-test bias. Unfortunately, as with other areas of epidemiology, intuition is often a poor statistician. In the 1990s, microbiologists adopted an intuitively appealing procedure referred to as discrepant analysis for the evaluation of new nucleic-acid-amplification tests (such as polymerase chain reaction) for the diagnosis of chlamydial infection and other infectious diseases. 4–6 Microbiologists recognized that these new tests represented a major advance over culture techniques, which were known to have limited sensitivity. 7 To address the concern of reference-test bias when using culture as the reference standard, the microbi-