Computer-aided diagnosis of renal obstruction: utility of log-linear modeling versus standard ROC and kappa analysis.

Computer-aided diagnosis of renal obstruction: utility of log-linear modeling versus standard ROC and kappa analysis.
复制标题

DOI:
10.1186/2191-219x-1-5
复制
发表时间:
2011-06-20
期刊:
影响因子:
3.2
通讯作者:
Taylor AT
Taylor AT
中科院分区:
医学3区
文献类型:
--
作者:
Manatunga AK;Binongo JN;Taylor AT

文献摘要

被引文献

相似文献

计算机辅助诊断 (CAD) 软件的准确性最好通过与代表疾病真实状态的黄金标准进行比较来评估。然而,在许多情况下,不可能了解疾病的真实状况,并且根据专家小组的解释来评估准确性。评估准确性的常见统计方法包括受试者工作特征 (ROC) 和 kappa 分析,但这两种方法都有很大的局限性,并且无法回答等效性问题:CAD 性能是否与专家的性能等效?本研究的目的是展示对数线性分析相对于标准 ROC 和 kappa 统计的优势,以评估计算机辅助诊断肾梗阻与专家读者提供的诊断的准确性。利用对数线性模型来分析先前发布的数据库,该数据库使用 ROC 和 kappa 统计数据来比较肾脏专家系统 (RENEX) 在 185 个肾脏(95 名患者)中生成的利尿肾图扫描解释(无梗阻、模棱两可或梗阻)与三位专家的独立和一致扫描解释,这三位专家对临床信息不知情,并前瞻性地独立地将每个肾脏分级为梗阻、模棱两可或非梗阻。对数线性模型表明,RENEX 和专家共识在无阻碍和有阻碍读数方面均具有超机会一致性(p < 0.0001)。此外,专家之间的成对一致性以及每位专家与 RENEX 之间的成对一致性没有显着差异(对于无阻碍、模棱两可和阻碍类别,p 分别 = 0.41、0.95、0.81)。同样,对于无阻碍 (p = 0.79) 和阻碍 (p = 0.49) 类别,三名专家的三向协议以及两名专家和 RENEX 的三向协议没有显着差异。对数线性模型表明,RENEX 相当于任何肾脏评级专家,特别是在梗阻和非梗阻类别方面。这个结论无法从原始的 ROC 和 kappa 分析中得出,它强调并说明了在没有黄金标准的情况下对数线性建模的作用和重要性。对数线性分析还提供了额外的证据,表明 RENEX 有潜力协助解释利尿肾图研究。
The accuracy of computer-aided diagnosis (CAD) software is best evaluated by comparison to a gold standard which represents the true status of disease. In many settings, however, knowledge of the true status of disease is not possible and accuracy is evaluated against the interpretations of an expert panel. Common statistical approaches to evaluate accuracy include receiver operating characteristic (ROC) and kappa analysis but both of these methods have significant limitations and cannot answer the question of equivalence: Is the CAD performance equivalent to that of an expert? The goal of this study is to show the strength of log-linear analysis over standard ROC and kappa statistics in evaluating the accuracy of computer-aided diagnosis of renal obstruction compared to the diagnosis provided by expert readers. Log-linear modeling was utilized to analyze a previously published database that used ROC and kappa statistics to compare diuresis renography scan interpretations (non-obstructed, equivocal, or obstructed) generated by a renal expert system (RENEX) in 185 kidneys (95 patients) with the independent and consensus scan interpretations of three experts who were blinded to clinical information and prospectively and independently graded each kidney as obstructed, equivocal, or non-obstructed. Log-linear modeling showed that RENEX and the expert consensus had beyond-chance agreement in both non-obstructed and obstructed readings (both p < 0.0001). Moreover, pairwise agreement between experts and pairwise agreement between each expert and RENEX were not significantly different (p = 0.41, 0.95, 0.81 for the non-obstructed, equivocal, and obstructed categories, respectively). Similarly, the three-way agreement of the three experts and three-way agreement of two experts and RENEX was not significantly different for non-obstructed (p = 0.79) and obstructed (p = 0.49) categories. Log-linear modeling showed that RENEX was equivalent to any expert in rating kidneys, particularly in the obstructed and non-obstructed categories. This conclusion, which could not be derived from the original ROC and kappa analysis, emphasizes and illustrates the role and importance of log-linear modeling in the absence of a gold standard. The log-linear analysis also provides additional evidence that RENEX has the potential to assist in the interpretation of diuresis renography studies.