Diminishing Uncertainty Within the Training Pool: Active Learning for Medical Image Segmentation

Diminishing Uncertainty Within the Training Pool: Active Learning for Medical Image Segmentation
复制标题

DOI:
10.1109/tmi.2020.3048055
复制
发表时间:
2021-10-01
影响因子:
10.6
通讯作者:
Roth, Holger R.
Roth, Holger R.
中科院分区:
工程技术1区
文献类型:
--
作者:
Nath, Vishwesh;Yang, Dong;Roth, Holger R.

文献摘要

被引文献

相似文献

主动学习是机器学习技术的独特抽象,其中模型/算法可以引导用户注释一组对模型有益的数据点,这与被动机器学习不同。主要优点是主动学习框架选择的数据点可以加速模型的学习过程,并且与在随机获取的数据集上训练的模型相比,可以减少实现完全准确性所需的数据量。已经提出了多种结合深度学习的主动学习框架,其中大多数都致力于分类任务。在这里,我们探讨主动学习的任务分割的医学成像数据集。我们使用两个数据集研究我们提出的框架:1。海马的MRI扫描,2.)胰腺和肿瘤的CT扫描。这项工作提出了一个查询由委员会的主动学习方法,其中联合优化器用于委员会。同时,我们提出了三个新的主动学习策略:1)。增加不确定数据的频率以偏置训练数据集; 2.)使用输入图像之间的互信息作为用于采集的正则化器,以确保训练数据集中的多样性; 3.)Stein变分梯度下降(SVGD)的Dice对数似然自适应。结果表明,通过实现完全准确性,同时仅使用每个数据集的22.69%和48.85%的可用数据,在数据减少方面有所改进。
Active learning is a unique abstraction of machine learning techniques where the model/algorithm could guide users for annotation of a set of data points that would be beneficial to the model, unlike passive machine learning. The primary advantage being that active learning frameworks select data points that can accelerate the learning process of a model and can reduce the amount of data needed to achieve full accuracy as compared to a model trained on a randomly acquired data set. Multiple frameworks for active learning combined with deep learning have been proposed, and the majority of them are dedicated to classification tasks. Herein, we explore active learning for the task of segmentation of medical imaging data sets. We investigate our proposed framework using two datasets: 1.) MRI scans of the hippocampus, 2.) CT scans of pancreas and tumors. This work presents a query-by-committee approach for active learning where a joint optimizer is used for the committee. At the same time, we propose three new strategies for active learning: 1.) increasing frequency of uncertain data to bias the training data set; 2.) Using mutual information among the input images as a regularizer for acquisition to ensure diversity in the training dataset; 3.) adaptation of Dice log-likelihood for Stein variational gradient descent (SVGD). The results indicate an improvement in terms of data reduction by achieving full accuracy while only using 22.69% and 48.85% of the available data for each dataset, respectively.