Modeling Perceptual Similarity Measures in CT Images of Focal Liver Lesions

Modeling Perceptual Similarity Measures in CT Images of Focal Liver Lesions
复制标题

DOI:
10.1007/s10278-012-9557-4
复制
发表时间:
2013-08-01
影响因子:
4.4
通讯作者:
Napel, Sandy
Napel, Sandy
中科院分区:
工程技术2区
文献类型:
--
作者:
Faruque, Jessica;Rubin, Daniel L.;Napel, Sandy

文献摘要

被引文献

相似文献

动机:医学图像感知相似性的黄金标准对于基于内容的图像检索至关重要,但读者之间的差异使开发变得复杂。我们的目标是开发一个统计模型,预测达到可接受的变异性水平所需的读者数量 (N)。材料和方法:我们收集了 3 位放射科医生对 171 对局灶性肝脏病变 CT 图像感知相似性的评分,按 9 分制评分。我们将读者的分数建模为加性高斯噪声中的双峰分布,并使用期望最大化算法根据分数估计分布参数。我们 (a) 采样了 171 个相似度分数来模拟真实情况,(b) 通过添加噪声来模拟读者,每个读者的标准差在 0 到 5 之间。我们计算了 2-50 个读者分数的平均值,并使用 Cohen 的 Kappa 度量计算了这些平均值与模拟地面事实之间的一致性 (AGT) 以及读者间一致性 (IRA)。结果:经验数据的 IRA 范围为 =0.41 至 0.66。对于 1.5 到 2.5 之间,三个模拟读者之间的 IRA 与经验数据中的一致性相当。对于这些值,AGT 的范围为 =0.81 至 0.91。正如预期的那样,AGT 随着 N 的增加而增加,当 N = 2 时,AGT 的范围分别为 0.83 至 0.92,当 N = 2 时,AGT 的范围分别为 50。结论:我们的模拟表明,对于中等至良好的 IRA,仍然可以获得出色的 AGT。该模型可用于预测所需的 N,以准确评估任意大小数据集中的相似性。
Motivation: A gold standard for perceptual similarity in medical images is vital to content-based image retrieval, but inter-reader variability complicates development. Our objective was to develop a statistical model that predicts the number of readers (N) necessary to achieve acceptable levels of variability. Materials and Methods: We collected 3 radiologists' ratings of the perceptual similarity of 171 pairs of CT images of focal liver lesions rated on a 9-point scale. We modeled the readers' scores as bimodal distributions in additive Gaussian noise and estimated the distribution parameters from the scores using an expectation maximization algorithm. We (a) sampled 171 similarity scores to simulate a ground truth and (b) simulated readers by adding noise, with standard deviation between 0 and 5 for each reader. We computed the mean values of 2-50 readers' scores and calculated the agreement (AGT) between these means and the simulated ground truth, and the inter-reader agreement (IRA), using Cohen's Kappa metric. Results: IRA for the empirical data ranged from =0.41 to 0.66. For between 1.5 and 2.5, IRA between three simulated readers was comparable to agreement in the empirical data. For these values, AGT ranged from =0.81 to 0.91. As expected, AGT increased with N, ranging from =0.83 to 0.92 for N = 2 to 50, respectively, with =2. Conclusion: Our simulations demonstrated that for moderate to good IRA, excellent AGT could nonetheless be obtained. This model may be used to predict the required N to accurately evaluate similarity in arbitrary size datasets.