AggNet: Deep Learning From Crowds for Mitosis Detection in Breast Cancer Histology Images

AggNet: Deep Learning From Crowds for Mitosis Detection in Breast Cancer Histology Images
复制标题

DOI:
10.1109/tmi.2016.2528120
复制
发表时间:
2016-05-01
影响因子:
10.6
通讯作者:
Navab, Nassir
Navab, Nassir
中科院分区:
工程技术1区
文献类型:
--
作者:
Albarqouni, Shadi;Baur, Christoph;Navab, Nassir

文献摘要

被引文献

相似文献

缺乏公开可用的地面实况数据已被确定为将深度学习的最新发展转移到生物医学成像领域的主要挑战。虽然众包已经实现了对真实的世界图像的大规模数据库的注释,但其用于生物医学目的的应用需要更深入的理解,因此需要对实际注释任务进行更精确的定义。专家任务被外包给非专家用户的事实可能会导致嘈杂的注释,从而在用户之间引入分歧。尽管传统的机器学习方法是从众包学习注释模型的宝贵资源,但在训练期间可能难以处理噪声注释。在本文中,我们提出了一个从人群中学习的新概念,通过额外的众包层(AggNet)直接处理数据聚合,作为卷积神经网络(CNN)学习过程的一部分。此外,我们提出了一个实验研究,旨在回答以下问题的群体学习。1)深度CNN可以用从众包收集的数据进行训练吗?2)如何调整CNN以在多种类型的注释数据集(地面实况和基于人群的)上进行训练?3)注释和聚合的选择如何影响准确性?我们的实验设置涉及Annot8,这是一个基于Crowdflower API的自实现网络平台,可为公开的生物医学图像数据库实现图像注释任务。我们的研究结果为从人群注释中进行深度CNN学习的功能提供了有价值的见解,并证明了数据聚合集成的必要性。
The lack of publicly available ground-truth data has been identified as the major challenge for transferring recent developments in deep learning to the biomedical imaging domain. Though crowdsourcing has enabled annotation of large scale databases for real world images, its application for biomedical purposes requires a deeper understanding and hence, more precise definition of the actual annotation task. The fact that expert tasks are being outsourced to non-expert users may lead to noisy annotations introducing disagreement between users. Despite being a valuable resource for learning annotation models from crowdsourcing, conventional machine-learning methods may have difficulties dealing with noisy annotations during training. In this manuscript, we present a new concept for learning from crowds that handle data aggregation directly as part of the learning process of the convolutional neural network (CNN) via additional crowdsourcing layer (AggNet). Besides, we present an experimental study on learning from crowds designed to answer the following questions. 1) Can deep CNN be trained with data collected from crowdsourcing? 2) How to adapt the CNN to train on multiple types of annotation datasets (ground truth and crowd-based)? 3) How does the choice of annotation and aggregation affect the accuracy? Our experimental setup involved Annot8, a self-implemented web-platform based on Crowdflower API realizing image annotation tasks for a publicly available biomedical image database. Our results give valuable insights into the functionality of deep CNN learning from crowd annotations and prove the necessity of data aggregation integration.