CHAOS Challenge- combined (CT-MR) healthy abdominal organ segmentation

CHAOS Challenge- combined (CT-MR) healthy abdominal organ segmentation
复制标题

DOI:
10.1016/j.media.2020.101950
复制
发表时间:
2021-01-06
影响因子:
10.9
通讯作者:
Selver, M. Alper
Selver, M. Alper
中科院分区:
工程技术1区
文献类型:
--
作者:
Kavur, A. Emre;Gezer, N. Sinem;Selver, M. Alper

文献摘要

被引文献

相似文献

多年来,腹部器官分割一直是一个综合性但尚未解决的研究领域。在过去的十年中,深度学习的密集发展引入了新的最先进的分割系统。尽管性能优于现有系统的总体精度,但动态链式模型的属性和参数对性能的影响很难解释。这使得比较分析成为通向解释性研究和系统的必要工具。此外,对于新兴的学习方法,如跨通道和多通道语义切分任务,DL的性能也很少被讨论。为了扩大关于这些主题的知识,与2019年在意大利威尼斯举行的IEEE国际生物医学成像研讨会(ISBI)一起组织了混沌组合(CT-MR)健康腹部器官分割挑战。从常规采集的腹部器官分割在几个临床应用中起着重要的作用,如手术前计划或各种疾病的形态和体积随访。这些应用程序需要在不同的指标集上达到一定水平的性能,例如最大对称表面距离(MSSD),以确定用于跟踪尺寸和形状差异的手术误差裕度或重叠误差。以前的腹部相关挑战主要集中在单一模式的肿瘤/病变检测和/或分类方面。相反,CHAOS提供了来自健康受试者的腹部CT和MR数据,用于单一和多个腹部器官的分割。设计了五项不同但相互补充的任务,以从多个角度分析参与方法的能力。通过与人工标注和交互方法的比较,对结果进行了深入的研究。分析表明,对于单通道(CT/MR),DL模型的性能可以表现出可靠的体积分析性能(骰子:0.98+/-0.00/0.95+/-0.01),但最好的MSSD性能仍然有限(21.89+/-13.94/20.85+/-10.63 mm)。对于肝脏的跨通道任务,参与模型的性能显著降低(骰子:0.88+/-0.15,MSSD:36.33+/-21.97 mm)。尽管在不同的应用上有相反的例子,但观察到旨在分割所有器官的多任务DL模型与特定于器官的模型相比表现更差(性能下降约5%)。尽管如此,一些成功的模型在多器官版本中表现得更好。我们的结论是,在单器官与多器官和跨器官分割中对这些利弊的探索将对开发支持真实世界临床应用的有效算法的进一步研究产生影响。最后,这项研究有1500多名参与者,收到了550多份意见书,这项研究的另一个重要贡献是分析了挑战组织的缺点,如多份意见书的影响和偷看现象。(C)2020爱思唯尔B.V.保留所有权利。
Segmentation of abdominal organs has been a comprehensive, yet unresolved, research field for many years. In the last decade, intensive developments in deep learning (DL) introduced new state-of-the-art segmentation systems. Despite outperforming the overall accuracy of existing systems, the effects of DL model properties and parameters on the performance are hard to interpret. This makes comparative anal-ysis a necessary tool towards interpretable studies and systems. Moreover, the performance of DL for emerging learning approaches such as cross-modality and multi-modal semantic segmentation tasks has been rarely discussed. In order to expand the knowledge on these topics, the CHAOS-Combined (CT-MR) Healthy Abdominal Organ Segmentation challenge was organized in conjunction with the IEEE Interna-tional Symposium on Biomedical Imaging (ISBI), 2019, in Venice, Italy. Abdominal organ segmentation from routine acquisitions plays an important role in several clinical applications, such as pre-surgical planning or morphological and volumetric follow-ups for various diseases. These applications require a certain level of performance on a diverse set of metrics such as maximum symmetric surface distance (MSSD) to determine surgical error-margin or overlap errors for tracking size and shape differences. Pre-vious abdomen related challenges are mainly focused on tumor/lesion detection and/or classification with a single modality. Conversely, CHAOS provides both abdominal CT and MR data from healthy subjects for single and multiple abdominal organ segmentation. Five different but complementary tasks were de-signed to analyze the capabilities of participating approaches from multiple perspectives. The results were investigated thoroughly, compared with manual annotations and interactive methods. The analysis shows that the performance of DL models for single modality (CT / MR) can show reliable volumetric analysis performance (DICE: 0.98 +/- 0.00 / 0.95 +/- 0.01), but the best MSSD performance remains limited (21.89 +/- 13.94 / 20.85 +/- 10.63 mm). The performances of participating models decrease dramatically for cross-modality tasks both for the liver (DICE: 0.88 +/- 0.15 MSSD: 36.33 +/- 21.97 mm). Despite contrary examples on different applications, multi-tasking DL models designed to segment all organs are observed to perform worse compared to organ-specific ones (performance drop around 5%). Nevertheless, some of the successful models show better performance with their multi-organ versions. We conclude that the exploration of those pros and cons in both single vs multi-organ and cross-modality segmentations is poised to have an impact on further research for developing effective algorithms that would support real-world clinical applications. Finally, having more than 1500 participants and receiving more than 550 submissions, another important contribution of this study is the analysis on shortcomings of challenge organizations such as the effects of multiple submissions and peeking phenomenon. (c) 2020 Elsevier B.V. All rights reserved.