Evaluation of uterine cervix segmentations using ground truth from multiple experts

Evaluation of uterine cervix segmentations using ground truth from multiple experts
复制标题

DOI:
10.1016/j.compmedimag.2008.12.002
复制
发表时间:
2009-04-01
影响因子:
5.7
通讯作者:
Greenspan, Hayit
Greenspan, Hayit
中科院分区:
工程技术2区
文献类型:
--
作者:
Gordon, Shiri;Lotenberg, Shelly;Greenspan, Hayit

文献摘要

被引文献

相似文献

这项工作的重点是为美国国家癌症研究所 (NCI) 收集的数字宫颈造影图像 (cervigrams) 的大型医学存储库生成和利用可靠的地面实况 (GT) 分割。 NCI 邀请了 20 名专家将一组 939 个宫颈图手动分割成医学和解剖学感兴趣的区域。基于这些独特的数据,当前工作的目标是:(1)自动生成多专家GT分割图; (2)使用GT图自动评估给定分割任务的复杂度; (3)使用GT图来评估自动分割算法的性能。多专家GT图是通过STAPLE(同时真理和性能水平估计)算法生成的,这是一种从多个观测值生成GT分割的众所周知的方法。定义了一种新的分割复杂性测量方法,该方法依赖于 GT 图内观察者间的变异性。该度量用于识别专家认为难以分割的图像,并比较不同分割任务的复杂性。提出了一种评估自动分割算法性能的准确性度量。使用所提出的精度测量来比较两种子宫颈边界检测算法。该度量反映了算法实现的实际分割质量。本文提出的方法和结论是通用的,可以应用于不同的图像和分割任务。在这里,它们被应用于宫颈图数据库,包括对可用数据的彻底分析。 (C) 2008 Elsevier Ltd. 保留所有权利。
This work is focused on the generation and utilization of a reliable ground truth (GT) segmentation for a large medical repository of digital cervicographic images (cervigrams) collected by the National Cancer Institute (NCI). NCI invited twenty experts to manually segment a set of 939 cervigrams into regions of medical and anatomical interest. Based on this unique data, the objectives of the current work are to: (1) Automatically generate a multi-expert GT segmentation map; (2) Use the GT map to automatically assess the complexity of a given segmentation task; (3) Use the GT map to evaluate the performance of an automated segmentation algorithm.The multi-expert GT map is generated via the STAPLE (Simultaneous Truth and Performance Level Estimation) algorithm, which is a well-known method to generate a GT segmentation from multiple observations. A new measure of segmentation complexity, which relies on the inter-observer variability within the GT map, is defined. This measure is used to identify images that were found difficult to segment by the experts and to compare the complexity of different segmentation tasks. An accuracy measure, which evaluates the performance of automated segmentation algorithms is presented. Two algorithms for cervix boundary detection are compared using the proposed accuracy measure. The measure is shown to reflect the actual segmentation quality achieved by the algorithms.The methods and conclusions presented in this work are general and can be applied to different images and segmentation tasks. Here they are applied to the cervigram database including a thorough analysis of the available data. (C) 2008 Elsevier Ltd. All rights reserved.