Why rankings of biomedical image analysis competitions should be interpreted with care

Why rankings of biomedical image analysis competitions should be interpreted with care
复制标题

DOI:
10.1038/s41467-018-07619-7
复制
发表时间:
2018-12-06
影响因子:
16.6
通讯作者:
Kopp-Schneider, Annette
Kopp-Schneider, Annette
中科院分区:
综合性期刊1区
文献类型:
--
作者:
Maier-Hein, Lena;Eisenmann, Matthias;Kopp-Schneider, Annette

文献摘要

被引文献

相似文献

国际挑战已成为验证生物医学图像分析方法的标准。鉴于他们的科学影响,令人惊讶的是,尚未对与挑战组织有关的共同实践进行批判性分析。在本文中,我们对迄今为止进行的生物医学图像分析挑战进行了全面分析。我们证明了挑战的重要性,并表明缺乏质量控制具有关键后果。首先,由于通常提供相关信息的一小部分,因此对结果的可重复性和解释通常受到阻碍。其次,对于许多变量,例如用于验证的测试数据,应用排名方案的测试数据和进行参考注释的观察者,算法的等级通常不鲁棒。为了克服这些问题,我们建议最佳实践指南,并定义将来要解决的开放研究问题。
International challenges have become the standard for validation of biomedical image analysis methods. Given their scientific impact, it is surprising that a critical analysis of common practices related to the organization of challenges has not yet been performed. In this paper, we present a comprehensive analysis of biomedical image analysis challenges conducted up to now. We demonstrate the importance of challenges and show that the lack of quality control has critical consequences. First, reproducibility and interpretation of the results is often hampered as only a fraction of relevant information is typically provided. Second, the rank of an algorithm is generally not robust to a number of variables such as the test data used for validation, the ranking scheme applied and the observers that make the reference annotations. To overcome these problems, we recommend best practice guidelines and define open research questions to be addressed in the future.