Optimization for Medical Image Segmentation: Theory and Practice When Evaluating With Dice Score or Jaccard Index

Optimization for Medical Image Segmentation: Theory and Practice When Evaluating With Dice Score or Jaccard Index
复制标题

DOI:
10.1109/tmi.2020.3002417
复制
发表时间:
2020-11-01
影响因子:
10.6
通讯作者:
Blaschko, Matthew B.
Blaschko, Matthew B.
中科院分区:
工程技术1区
文献类型:
--
作者:
Eelbode, Tom;Bertels, Jeroen;Blaschko, Matthew B.

文献摘要

被引文献

相似文献

在许多医学成像和经典的计算机视觉任务中,Dice分数和Jaccard指数被用来评估分割性能。尽管度量敏感损失的存在和实验取得了巨大的成功,即放松了这些度量,如Soft Dice、Soft Jaccard和Lovasz-Softmax,但许多研究人员仍然使用每像素损失,如(加权的)交叉熵来训练CNN用于分割。因此,目标指标在许多情况下没有直接优化。我们从理论的角度研究了度量敏感损失函数组内的关系,并质疑是否存在加权交叉熵的最优加权方案来优化测试时的Dice分数和Jaccard指数。我们发现Dice分数和Jaccard指数是相对和绝对接近的,但对于加权Hamming相似,我们没有发现这样的近似。对于Tversky损失,当偏离柔和Tversky等于软骰子的微不足道的权重设置时,近似变得单调更差。我们在对六个医学分割任务的广泛验证中经验地验证了这些结果,并且可以证实,在使用Dice Score或Jaccard Index进行评估的情况下,度量敏感损失优于基于交叉熵的损失函数。这进一步适用于多类设置,并且跨越不同的对象大小和前景/背景比。这些结果鼓励在医疗细分任务中更广泛地采用度量敏感的损失函数,其中感兴趣的性能衡量标准是Dice分数或Jaccard指数。
In many medical imaging and classical computer vision tasks, the Dice score and Jaccard index are used to evaluate the segmentation performance. Despite the existence and great empirical success of metric-sensitive losses, i.e. relaxations of these metrics such as soft Dice, soft Jaccard and Lovasz-Softmax, many researchers still use per-pixel losses, such as (weighted) cross-entropy to train CNNs for segmentation. Therefore, the target metric is in many cases not directly optimized. We investigate from a theoretical perspective, the relation within the group of metric-sensitive loss functions and question the existence of an optimal weighting scheme for weighted cross-entropy to optimize the Dice score and Jaccard index at test time. We find that the Dice score and Jaccard index approximate each other relatively and absolutely, but we find no such approximation for a weighted Hamming similarity. For the Tversky loss, the approximation gets monotonically worse when deviating from the trivial weight setting where soft Tversky equals soft Dice. We verify these results empirically in an extensive validation on six medical segmentation tasks and can confirm that metric-sensitive losses are superior to cross-entropy based loss functions in case of evaluation with Dice Score or Jaccard Index. This further holds in a multi-class setting, and across different object sizes and foreground/background ratios. These results encourage a wider adoption of metric-sensitive loss functions for medical segmentation tasks where the performance measure of interest is the Dice score or Jaccard index.