An active learning based classification strategy for the minority class problem: application to histopathology annotation.

An active learning based classification strategy for the minority class problem: application to histopathology annotation.
复制标题

DOI:
10.1186/1471-2105-12-424
复制
发表时间:
2011-10-28
期刊:
影响因子:
3
通讯作者:
Madabhushi A
Madabhushi A
中科院分区:
生物学4区
文献类型:
--
作者:
Doyle S;Monaco J;Feldman M;Tomaszewski J;Madabhushi A

文献摘要

参考文献

被引文献

相似文献

数字病理学的监督分类器可以提高医生检测和诊断癌症等疾病的能力。为分类器生成训练数据是有问题的,因为只有领域专家(例如病理学家)才能正确标记地面实况数据。此外,数字病理学数据集遭受“少数类问题”,即来自非目标类的样本数量超过目标类样本的数量的问题,这可能使分类器产生偏差并降低准确性。在本文中,我们开发了一种结合主动学习(AL)和类平衡的训练策略。AL识别具有“信息性”(即可能提高分类器性能)的未标记样本进行注释,避免无信息样本。与随机学习(RL)相比,这可以在较小的训练集大小下获得高精度。以前的AL方法没有明确占少数类问题的生物医学图像。预先指定目标类别比例可以缓解训练偏差的问题。最后,我们开发了一个数学模型来预测实现均衡训练类所需的注释数量(成本)。除了预测培训成本,该模型揭示了在少数类问题的背景下AL的理论属性。使用这种类平衡的AL训练策略(CBAL),我们建立了一个分类器,以区分数字化前列腺组织病理学上的癌症和非癌症区域。我们的数据集包括从100个活检样本(58名前列腺癌患者)中采样的12,000个图像区域。我们将CBAL与以下各项进行比较:(1)不平衡AL(UBAL),使用AL但忽略类别比率;(2)类别平衡RL(CBRL),使用具有特定类别比率的RL;以及(3)不平衡RL(UBRL)。CBAL训练的分类器比交替训练的分类器产生高2%的准确性和高3%的受试者工作特征曲线(AUC)下的面积。我们的成本模型准确地预测了获得平衡类所需的注释数量。我们的预测的准确性验证了实验观察到的成本。最后,我们发现,过采样的少数类产生的分类精度的边际改善,但改善的性能是以牺牲更大的注释成本。我们将AL与类平衡相结合,以产生适用于大多数监督分类问题的一般训练策略,其中数据集获取成本很高,并且存在少数类问题。智能训练策略是监督分类的关键组成部分,但AL和类比的智能选择的集成,以及通用成本模型的应用,将有助于研究人员更快,更有效地规划训练过程。
Supervised classifiers for digital pathology can improve the ability of physicians to detect and diagnose diseases such as cancer. Generating training data for classifiers is problematic, since only domain experts (e.g. pathologists) can correctly label ground truth data. Additionally, digital pathology datasets suffer from the "minority class problem", an issue where the number of exemplars from the non-target class outnumber target class exemplars which can bias the classifier and reduce accuracy. In this paper, we develop a training strategy combining active learning (AL) with class-balancing. AL identifies unlabeled samples that are "informative" (i.e. likely to increase classifier performance) for annotation, avoiding non-informative samples. This yields high accuracy with a smaller training set size compared with random learning (RL). Previous AL methods have not explicitly accounted for the minority class problem in biomedical images. Pre-specifying a target class ratio mitigates the problem of training bias. Finally, we develop a mathematical model to predict the number of annotations (cost) required to achieve balanced training classes. In addition to predicting training cost, the model reveals the theoretical properties of AL in the context of the minority class problem. Using this class-balanced AL training strategy (CBAL), we build a classifier to distinguish cancer from non-cancer regions on digitized prostate histopathology. Our dataset consists of 12,000 image regions sampled from 100 biopsies (58 prostate cancer patients). We compare CBAL against: (1) unbalanced AL (UBAL), which uses AL but ignores class ratio; (2) class-balanced RL (CBRL), which uses RL with a specific class ratio; and (3) unbalanced RL (UBRL). The CBAL-trained classifier yields 2% greater accuracy and 3% higher area under the receiver operating characteristic curve (AUC) than alternatively-trained classifiers. Our cost model accurately predicts the number of annotations necessary to obtain balanced classes. The accuracy of our prediction is verified by empirically-observed costs. Finally, we find that over-sampling the minority class yields a marginal improvement in classifier accuracy but the improved performance comes at the expense of greater annotation cost. We have combined AL with class balancing to yield a general training strategy applicable to most supervised classification problems where the dataset is expensive to obtain and which suffers from the minority class problem. An intelligent training strategy is a critical component of supervised classification, but the integration of AL and intelligent choice of class ratios, as well as the application of a general cost model, will help researchers to plan the training process more quickly and effectively.
DOI: 10.1109/tpami.2006.156
发表时间: 2006-08-01
影响因子: 23.6
作者:
Li, Mingkun;Sethi, Ishwar K.
通讯作者: Sethi, Ishwar K.
DOI: 10.1109/tcom.1983.1095851
发表时间: 1983-01-01
影响因子: 8.3
作者:
BURT, PJ;ADELSON, EH
通讯作者: ADELSON, EH
DOI: 10.1109/34.531803
发表时间: 1996-08-01
影响因子: 23.6
作者:
Manjunath, BS;Ma, WY
通讯作者: Ma, WY
DOI: 10.1109/tsmc.1973.4309314
发表时间: 1973-01-01
期刊: IEEE TRANSACTIONS ON SYSTEMS MAN AND CYBERNETICS
影响因子: --
作者:
HARALICK, RM;SHANMUGAM, K;DINSTEIN, I
通讯作者: DINSTEIN, I
DOI: 10.1613/jair.953
发表时间: 2002-01-01
影响因子: 5
作者:
Chawla, NV;Bowyer, KW;Kegelmeyer, WP
通讯作者: Kegelmeyer, WP