Robust statistical label fusion through COnsensus Level, Labeler Accuracy, and Truth Estimation (COLLATE).

Robust statistical label fusion through COnsensus Level, Labeler Accuracy, and Truth Estimation (COLLATE).
复制标题

DOI:
10.1109/tmi.2011.2147795
复制
发表时间:
2011-10
影响因子:
10.6
通讯作者:
Landman BA
Landman BA
中科院分区:
工程技术1区
文献类型:
--
作者:
Asman AJ;Landman BA

文献摘要

被引文献

相似文献

医学图像中感兴趣结构的分割和描绘对于量化和描述与临床相关条件的结构、形态和功能的相关性至关重要。已建立的分割的黄金标准是由神经解剖学家专家手动逐个体素标记。这一过程可能非常耗时、资源密集,而且观察员之间的变异性很大。因此,涉及新奇结构或外观特征的研究在范围(受试者数量)、规模(评估的区域范围)和统计能力方面受到限制。已经提出了融合来自几个不同来源(例如,多个人类观察者)的数据集的统计方法,以同时估计评分者的表现和地面事实标签。然而,对于经验数据集,统计融合已被观察到导致视觉上不一致的结果。因此,尽管统计方法简单易行,但在实践中经常使用单一观察员和/或直接投票。因此,在标签估计过程中,评分员的表现没有被系统地量化和利用。到目前为止,统计融合方法一直依赖于评分员表现的特征,而这些特征本质上并不包括评分员表现的空间变化模型。在这里,我们提出了一种新的、稳健的统计标签融合算法来估计和解释空间变化的性能。该算法(Consensus Level,Labeler Accuracy and Truth Estiment(COLLATE))基于一个简单的概念,即图像的一些区域很难标记(例如,混淆区域:边界或低对比度区域),而其他区域本质上是显而易见的(例如,共识区域:大区域或高对比度边缘的中心)。与它的前身不同,COLLATE估计每个体素的共识水平,并估计每个地区不同的观察者行为模型。我们表明,在模拟和经验数据集中,COLLATE在标签准确性和评分者评估方面都比以前的融合方法有了显着的改进。
Segmentation and delineation of structures of interest in medical images is paramount to quantifying and characterizing structural, morphological, and functional correlations with clinically relevant conditions. The established gold standard for performing segmentation has been manual voxel-by-voxel labeling by a neuroanatomist expert. This process can be extremely time consuming, resource intensive and fraught with high inter-observer variability. Hence, studies involving characterizations of novel structures or appearances have been limited in scope (numbers of subjects), scale (extent of regions assessed), and statistical power. Statistical methods to fuse data sets from several different sources (e.g., multiple human observers) have been proposed to simultaneously estimate both rater performance and the ground truth labels. However, with empirical datasets, statistical fusion has been observed to result in visually inconsistent findings. So, despite the ease and elegance of a statistical approach, single observers and/or direct voting are often used in practice. Hence, rater performance is not systematically quantified and exploited during label estimation. To date, statistical fusion methods have relied on characterizations of rater performance that do not intrinsically include spatially varying models of rater performance. Herein, we present a novel, robust statistical label fusion algorithm to estimate and account for spatially varying performance. This algorithm, COnsensus Level, Labeler Accuracy and Truth Estimation (COLLATE), is based on the simple idea that some regions of an image are difficult to label (e.g., confusion regions: boundaries or low contrast areas) while other regions are intrinsically obvious (e.g., consensus regions: centers of large regions or high contrast edges). Unlike its predecessors, COLLATE estimates the consensus level of each voxel and estimates differing models of observer behavior in each region. We show that COLLATE provides significant improvement in label accuracy and rater assessment over previous fusion methods in both simulated and empirical datasets.