Assessing operating characteristics of CAD algorithms in the absence of a gold standard.

Assessing operating characteristics of CAD algorithms in the absence of a gold standard.
复制标题

在缺乏黄金标准的情况下评估 CAD 算法的操作特性。

DOI:
10.1118/1.3352687
复制
发表时间:
2010
期刊:
影响因子:
3.8
通讯作者:
Rubin,GeoffreyD
Rubin,GeoffreyD
中科院分区:
医学3区
文献类型:
--
作者:
Choudhury,KingshukRoy;Paik,DavidS;Yi,ChinA;Napel,Sandy;Roos,Justus;Rubin,GeoffreyD

文献摘要

被引文献

相似文献

目的:作者在使用参考阅读器面板作为估计CAD算法检测病变的操作特性的“金标准”时,检查潜在的偏差。作为一种替代方案,作者提出了潜在类分析(LCA),它不需要外部金标准来评估诊断的准确性。方法建立不同诊断方案下的多读卡器检测二项模型,假设读卡器在真实病变状态下具有条件独立性。利用最大似然LCA估计各方案的工作特性。使用从二项模型中模拟的数据对一系列操作特性进行比较,阅读器面板和基于LCA的估计。LCA应用于来自肺图像数据库联盟(LIDC)的36个薄片胸部计算机断层扫描数据集:将4名放射科医生的自由搜索标记与4名不同CAD辅助放射科医生的标记进行比较。对于真实数据,提出了基于自举的重采样方法,该方法适应阅读器检测中的依赖性,以检验检测协议之间差异的假设。结果在模拟研究中,基于读者面板的灵敏度估计的平均相对偏差(ARB)为- 23%至- 27%,显著高于LCA (ARB为- 2%至- 6%)。通过读者组(ARB为- 0.6% - 0.5%)和LCA (ARB为1.4%-0.5%)可以很好地估计特异性。在LIDC考虑的1145个候选病变中,LCA估计参考阅读器的敏感度(55%)显著低于CAD辅助阅读器(68%)(‐value 0.006)。参考阅读器的每位患者平均假阳性(0.95)并不显著低于CAD辅助阅读器(1.27)(‐值0.28)。结论:虽然基于读者共识的金标准可能会严重偏倚敏感性估计,但LCA可能是评估诊断准确性的更准确和一致的方法。
PurposeThe authors examine potential bias when using a reference reader panel as “gold standard” for estimating operating characteristics of CAD algorithms for detecting lesions. As an alternative, the authors propose latent class analysis (LCA), which does not require an external gold standard to evaluate diagnostic accuracy.MethodsA binomial model for multiple reader detections using different diagnostic protocols was constructed, assuming conditional independence of readings given true lesion status. Operating characteristics of all protocols were estimated by maximum likelihood LCA. Reader panel and LCA based estimates were compared using data simulated from the binomial model for a range of operating characteristics. LCA was applied to 36 thin section thoracic computed tomography data sets from the Lung Image Database Consortium (LIDC): Free search markings of four radiologists were compared to markings from four different CAD assisted radiologists. For real data, bootstrap‐based resampling methods, which accommodate dependence in reader detections, are proposed to test of hypotheses of differences between detection protocols.ResultsIn simulation studies, reader panel based sensitivity estimates had an average relative bias (ARB) of −23% to −27%, significantly higher (‐value ) than LCA (ARB −2% to −6%). Specificity was well estimated by both reader panel (ARB −0.6% to −0.5%) and LCA (ARB 1.4%–0.5%). Among 1145 lesion candidates LIDC considered, LCA estimated sensitivity of reference readers (55%) was significantly lower (‐value 0.006) than CAD assisted readers’ (68%). Average false positives per patient for reference readers (0.95) was not significantly lower (‐value 0.28) than CAD assisted readers’ (1.27).ConclusionsWhereas a gold standard based on a consensus of readers may substantially bias sensitivity estimates, LCA may be a significantly more accurate and consistent means for evaluating diagnostic accuracy.