Estimating latent reader-performance variability using the Obuchowski-Rockette method.

Estimating latent reader-performance variability using the Obuchowski-Rockette method.
复制标题

使用 Obuchowski-Rockette 方法估计潜在的读者性能变异性。

DOI:
10.1117/12.2513106
复制
发表时间:
2019
期刊:
Proceedings of SPIE--the International Society for Optical Engineering
影响因子:
--
通讯作者:
Brennan,PatrickC
Brennan,PatrickC
中科院分区:
--
文献类型:
--
作者:
Hillis,StephenL;Mohammad,BaderaAl;Brennan,PatrickC

文献摘要

相似文献

我们描述了如何使用多读者诊断研究的Obuchowski-Rockette(OR)分析方法来估计潜在读者表现结果的变异性,例如ROC曲线下面积(AUC)。对于一个特定的读者,潜在的或真实的读者绩效结果可以在概念上被认为是如果读者阅读大量案例的估计结果。我们注意到,对于典型诊断研究中使用的样本量,潜在阅片者表现结果等于观察结果减去测量误差。一项经常被引用的研究评估了各种阅片者表现结果的可变性,包括AUC,这是克雷格比姆等人的研究。例如,“美国放射科医师对筛查乳房X线照片解释的变异性”,1996年出版。然而,这种类型的研究的一个问题是,变异性估计包括测量误差。因此,这种方法高估了潜在的读者变异性,并给出了依赖于案例样本大小的变异性估计。所提出的方法克服了这些问题。我们说明了在约旦的29个放射科医生,每个阅读60胸部计算机断层扫描(CT)扫描的建议方法。使用OR方法,我们能够估计潜伏AUC值的中间95%范围为0.07;即,我们估计95%的放射科医师在他们成功区分一对患病和非患病病例的能力上相差小于0.07。相比之下,观察到的AUC的95%范围的估计值为0.18。因此,我们可以看到,描述读者可变性的传统方法,会大大夸大读者真实能力的可变性。
We describe how the Obuchowski-Rockette (OR) method of analysis for multi-reader diagnostic studies can be used to estimate the variability of latent reader-performance outcomes, such as the area under the ROC curve (AUC). For a specific reader the latent or true reader performance outcome can conceptually be thought of as the estimate that would result if the reader were to read a very large number of cases. We note that for the sample sizes used in typical diagnostic studies, the latent reader-performance outcome is equal to the observed outcome minus measurement error. An often-cited study that assesses the variability of various reader-performance outcomes, including the AUC, is the study by Craig Beam et. al., “Variability in the Interpretation of Screening Mammograms by US Radiologists,” published in 1996. However, a problem with this type of study is that the variability estimates includes measurement error. Thus this approach overestimates latent reader variability and gives variability estimates that are dependent on case sample size. The proposed method overcomes these problems. We illustrate the proposed method for 29 radiologists in Jordan, with each reading 60 chest computed tomography (CT) scans. Using the OR method we were able to estimate the middle 95% range for latent AUC values to be 0.07; i.e., we estimate that 95% of radiologists differ by less than 0.07 in their ability to successfully discriminate between a pair of diseased and non-diseased cases. In contrast, the estimate for the 95% range for the observed AUCs was 0.18. Thus we see how conventional methods of describing reader variability can greatly overstate the variability of the true abilities of the readers.