Measuring rater bias in diagnostic tests with ordinal ratings.

Measuring rater bias in diagnostic tests with ordinal ratings.
复制标题

DOI:
10.1002/sim.9011
复制
发表时间:
2021-07-30
影响因子:
2
通讯作者:
Nelson KP
Nelson KP
中科院分区:
医学3区
文献类型:
--
作者:
Kim C;Lin X;Nelson KP

文献摘要

参考文献

相似文献

诊断测试通常依赖于熟练的评估者对图像的解释。然而,在许多临床环境中,专家评级之间观察到的差异会对这些解释的置信度产生不利影响,从而导致诊断过程的不确定性。例如,在乳腺癌检测中,放射科医生解释乳房X线照相图像,而乳腺活检结果由病理学家检查。这些程序中的每一个都涉及主观性因素。我们在这里提出了一种灵活的两阶段贝叶斯潜变量模型,以研究个体评估者的技能如何影响大规模医学检测研究中图像相关检测的诊断准确性。所提出模型的优点在于,患者在合理时间范围内的真实疾病状态可能是已知的,也可能是未知的。在这些研究中,许多评估者各自使用定义的顺序分级量表对大量患者样本进行分类,从而导致评级之间存在复杂的相关结构。与当前可用的方法相比,我们的建模方法考虑了专家和患者贡献的不同变异来源,同时考虑了评级和患者之间存在的相关性。我们提出了一种新的评估者能力衡量标准(放大镜),与传统的敏感性和特异性衡量标准相比,它对人群中疾病的潜在患病率具有稳健性,为整个患者群体的诊断准确性提供了另一种衡量标准。广泛的模拟研究表明,参数估计和准确性测量的偏差较低,并说明所提出的模型与现有模型相比具有更好的性能。导出受试者工作特征(ROC)曲线以评估个别专家的诊断准确性及其整体表现。我们提出的建模方法适用于已知疾病状态的大型乳腺成像研究和未知疾病状态的子宫癌数据集。
Diagnostic tests are frequently reliant upon the interpretation of images by skilled raters. In many clinical settings, however, the variability observed between experts’ ratings plays a detrimental role in the degree of confidence in these interpretations, leading to uncertainty in the diagnostic process. For example, in breast cancer testing, radiologists interpret mammographic images, while breast biopsy results are examined by pathologists. Each of these procedures involves elements of subjectivity. We propose here a flexible two-stage Bayesian latent variable model to investigate how the skills of individual raters impact the diagnostic accuracy of image-related testing in large-scale medical testing studies. A strength of the proposed model is that the true disease status of a patient within a reasonable time frame may or may not be known. In these studies, many raters each contribute classifications on a large sample of patients using a defined ordinal grading scale, leading to a complex correlation structure between ratings. Our modeling approach considers the different sources of variability contributed by experts and patients while accounting for correlations present between ratings and patients, in contrast to currently available methods. We propose a novel measure of a rater’s ability (magnifier) that, in contrast to conventional measures of sensitivity and specificity, is robust to the underlying prevalence of disease in the population, providing an alternative measure of diagnostic accuracy across patient populations. Extensive simulation studies demonstrate lower bias in estimation of parameters and measures of accuracy, and illustrate outperformance of the proposed model when compared to existing models. Receiver operator characteristic (ROC) curves are derived to assess the diagnostic accuracy of individual experts and their overall performance. Our proposed modeling approach is applied to a large breast imaging study for known disease status and a uterine cancer dataset for unknown disease status.
DOI: 10.1198/106186005x63185
发表时间: 2005-09-01
影响因子: 2.4
作者:
Kottas, A;Müller, P;Quintana, F
通讯作者: Quintana, F
DOI: 10.1002/cjs.11253
发表时间: 2015-09-01
影响因子: 0.6
作者:
Bao, Junshu;Hanson, Timothy E.
通讯作者: Hanson, Timothy E.
DOI: 10.1080/10543406.2016.1226334
发表时间: 2016-01-01
影响因子: 1.1
作者:
Collins, John;Albert, Paul S.
通讯作者: Albert, Paul S.
DOI: 10.3102/1076998609353116
发表时间: 2010-04-01
影响因子: 2.4
作者:
Cao, Jing;Stokes, S. Lynne;Zhang, Song
通讯作者: Zhang, Song
DOI: 10.1093/jnci/95.4.282
发表时间: 2003-02-19
期刊: JOURNAL OF THE NATIONAL CANCER INSTITUTE
影响因子: --
作者:
Beam, CA;Conant, EF;Sickles, EA
通讯作者: Sickles, EA