Simultaneous truth and performance level estimation (STAPLE): An algorithm for the validation of image segmentation

Simultaneous truth and performance level estimation (STAPLE): An algorithm for the validation of image segmentation
复制标题

DOI:
10.1109/tmi.2004.828354
复制
发表时间:
2004-07-01
影响因子:
10.6
通讯作者:
Wells, WM
Wells, WM
中科院分区:
工程技术1区
文献类型:
--
作者:
Warfield, SK;Zou, KH;Wells, WM

文献摘要

被引文献

相似文献

表征图像分割方法的性能一直是一个持久的挑战。由于分割算法的精确度和精确度往往有限,因此性能分析非常重要。由人类评分员交互绘制所需的细分往往是唯一可接受的方法,但存在评分者内部和评分者之间的可变性。为了消除评分者引入的变异性,已经寻求了自动算法,但必须对这种算法进行评估,以确保它们适合于这项任务。由于难以获得或估计临床数据的已知真实分割,评分者(人工或算法)生成医学图像分割的性能一直难以量化。虽然可以构建已知或容易估计其地面真相的物理和数字体模,但由于构建再现临床数据中观察到的全部成像特征以及正常和病理解剖学变异性的体模的难度,这些体模不能完全反映临床图像。与评分者的分割集合进行比较是一个有吸引力的替代方案,因为它可以直接在相关的临床影像数据上进行。然而,目前还没有明确的最合适的度量或度量集来比较这些分割,并且在实践中使用了几个度量。本文提出了一种同时进行真实性和性能水平估计的期望最大化算法。该算法考虑分割的集合,并计算真实分割的概率估计和由每个分割表示的性能水平的测量。集合中每个分割的源可以是经过适当训练的一个或多个人类评分者,或者可以是自动分割算法。真实分割的概率估计是通过估计分割的最佳组合、根据估计的性能水平对每个分割进行加权、以及结合用于被分割的结构的空间分布的先验模型以及空间同质性约束来形成的。史泰博可以直接应用于临床影像数据,它可以方便地评估自动图像分割算法的性能,并可以直接比较人类评分员和算法的性能。
Characterizing the performance of image segmentation approaches has been a persistent challenge. Performance analysis is important since segmentation algorithms often have limited accuracy and precision. Interactive drawing of the desired segmentation by human raters has often been the only acceptable approach, and yet suffers from intra-rater and inter-rater variability. Automated algorithms have been sought in order to remove the variability introduced by raters, but such algorithms must be assessed to ensure they are suitable for the task.The performance of raters (human or algorithmic) generating segmentations of medical images has been difficult to quantify because of the difficulty of obtaining or estimating a known true segmentation for clinical data. Although physical and digital phantoms can be constructed for which ground truth is known or readily estimated, such phantoms do not fully reflect clinical images due to the difficulty of constructing phantoms which reproduce the full range of imaging characteristics and normal and pathological anatomical variability observed in clinical data. Comparison to a collection of segmentations by raters is an attractive alternative since it can be carried out directly on the relevant clinical imaging data. However, the most appropriate measure or set of measures with which to compare such segmentations has not been clarified and several measures are used in practice.We present here an expectation-maximization algorithm for simultaneous truth and performance level estimation (STAPLE). The algorithm considers a collection of segmentations and computes a probabilistic estimate of the true segmentation and a measure of the performance level represented by each segmentation. The source of each segmentation in the collection may be an appropriately trained human rater or raters, or may be an automated segmentation algorithm. The probabilistic estimate of the true segmentation is formed by estimating an optimal combination of the segmentations, weighting each segmentation depending upon the estimated performance level, and incorporating a prior model for the spatial distribution of structures being segmented as well as spatial homogeneity constraints. STAPLE is straightforward to apply to clinical imaging data, it readily enables assessment of the performance of an automated image segmentation algorithm, and enables direct comparison of human rater and algorithm performance.