Analysis of Underlying Causes of Inter-expert Disagreement in Retinopathy of Prematurity Diagnosis Application of Machine Learning Principles

Analysis of Underlying Causes of Inter-expert Disagreement in Retinopathy of Prematurity Diagnosis Application of Machine Learning Principles
复制标题

DOI:
10.3414/me13-01-0081
复制
发表时间:
2015-01-01
影响因子:
1.7
通讯作者:
Chiang, M. F.
Chiang, M. F.
中科院分区:
医学4区
文献类型:
--
作者:
Ataer-Cansizoglu, E.;Kalpathy-Cramer, J.;Chiang, M. F.

文献摘要

被引文献

相似文献

目的:基于图像的临床诊断的专家间差异已在许多疾病中得到证实,包括早产儿视网膜病变(ROP),这是一种影响低出生体重婴儿的疾病,也是儿童失明的主要原因。为了更好地理解专家之间变异性的根本原因,我们提出了一种量化专家决策变异性并分析专家诊断与图像计算特征之间关系的方法。这些特征的识别与 ROP 中基于计算机的决策支持系统和教育系统的开发相关,并且这些方法可能适用于观察到专家间变异的其他疾病。 方法:实验在 34 个视网膜图像的数据集上进行,每个图像都有 22 名专家独立提供的诊断。使用互信息 (MI) 和核密度估计的概念进行分析。从视网膜图像中提取了大量的结构特征(总共 66 个)。 22 位研究专家利用特征选择来识别与实际临床决策相关的最重要特征。通过对所有可能的特征子集进行详尽的搜索并考虑联合 MI 作为相关性标准,为每个观察者选择最好的三个特征。我们还将我们的结果与 Cohen 的 Kappa [36] 的结果进行了比较,作为评估者间的可靠性度量。结果:结果表明,一组观察者(22 人中的 17 人)做出的决定彼此一致。小动脉迂曲度的平均和第二中心矩是该小组与其他观察者之间意见不一致的原因之一,这意味着专家组考虑了迂曲度的数量以及图像中迂曲度的变化。结论:给定一组基于图像的特征,所提出的分析方法可以识别导致专家在 ROP 诊断中达成一致和分歧的关键基于图像的特征。尽管基于树的特征和各种统计数据(例如中心矩)在文献中并不流行,但我们的结果表明它们对于诊断很重要。
Objective: Inter-expert variability in image-based clinical diagnosis has been demonstrated in many diseases including retinopathy of prematurity (ROP), which is a disease affecting low birth weight infants and is a major cause of childhood blindness. In order to better understand the underlying causes of variability among experts, we propose a method to quantify the variability of expert decisions and analyze the relationship between expert diagnoses and features computed from the images. Identification of these features is relevant for development of computer-based decision support systems and educational systems in ROP, and these methods may be applicable to other diseases where inter-expert variability is observed.Methods: The experiments were carried out on a dataset of 34 retinal images, each with diagnoses provided independently by 22 experts. Analysis was performed using concepts of Mutual Information (MI) and Kernel Density Estimation. A large set of structural features (a total of 66) were extracted from retinal images. Feature selection was utilized to identify the most important features that correlated to actual clinical decisions by the 22 study experts. The best three features for each observer were selected by an exhaustive search on all possible feature subsets and considering joint MI as a relevance criterion. We also compared our results with the results of Cohen's Kappa [36] as an inter-rater reliability measure.Results: The results demonstrate that a group of observers (17 among 22) decide consistently with each other. Mean and second central moment of arteriolar tortuosity is among the reasons of disagreement between this group and the rest of the observers, meaning that the group of experts consider amount of tortuosity as well as the variation of tortuosity in the image.Conclusion: Given a set of image-based features, the proposed analysis method can identify critical image-based features that lead to expert agreement and disagreement in diagnosis of ROP. Although tree-based features and various statistics such as central moment are not popular in the literature, our results suggest that they are important for diagnosis.