A new co-training-style random forest for computer aided diagnosis

A new co-training-style random forest for computer aided diagnosis
复制标题

用于计算机辅助诊断的新型协同训练式随机森林

DOI:
10.1007/s10844-009-0105-8
复制
发表时间:
2011-06
影响因子:
3.4
通讯作者:
M. Zu Guo
M. Zu Guo
中科院分区:
计算机科学3区
文献类型:
--
作者:
Chao Deng;M. Zu Guo

文献摘要

参考文献

被引文献

相似文献

计算机辅助诊断(CAD)系统中使用的机器学习技术学习一个假设,以帮助医学专家在未来做出诊断。为了学习一个性能良好的假设,需要大量的专家诊断的例子,这给专家带来了沉重的负担。协同训练式随机森林(Co-Forest)利用大量未诊断样本和集成学习的能力,减轻了专家的负担,产生性能良好的假设。然而,Co-forest可能会遇到其他co-training风格算法所共有的问题,即未标记的示例可能是在训练过程中积累的错误标记的示例。这是由于有限数量的原始标记的例子通常产生差的组件分类器,缺乏多样性和准确性。本文提出了一种新的Co-Forest算法--具有自适应数据编辑的Co-Forest(ADE-Co-Forest)。它不仅利用了一种特定的数据编辑技术,以识别和丢弃可能被错误标记的例子在整个共同标记迭代,但它也采用了自适应策略,以决定是否触发编辑操作,根据不同的情况。自适应策略结合了五个预条件定理,所有这些都确保了迭代减少分类错误和PAC学习理论下的新的训练集的规模增加。在UCI数据集上的实验和使用胸部CT图像进行肺小结节检测的应用表明,ADE-Co-Forest比Co-Forest和DE-Co-Forest(具有数据编辑但没有自适应策略的Co-Forest)更有效地提高了学习假设的性能。
Machine learning techniques used incomputer aided diagnosis(CAD) systems learn a hypothesis to help the medical experts make a diagnosis in the future. To learn a well-performed hypothesis, a large amount of expert-diagnosed examples are required, which places a heavy burden on experts. By exploiting large amounts of undiagnosed examples and the power of ensemble learning, theco-training-style random forest(Co-Forest) releases the burden on the experts and produces well-performed hypotheses. However, the Co-forest may suffer from a problem common to other co-training-style algorithms, namely, that the unlabeled examples may instead be wrongly-labeled examples that become accumulated in the training process. This is due to the fact that the limited number of originally-labeled examples usually produces poor component classifiers, which lack diversity and accuracy. In this paper, a new Co-Forest algorithm namedCo-Forest with Adaptive Data Editing(ADE-Co-Forest) is proposed. Not only does it exploit a specific data-editing technique in order to identify and discard possibly mislabeled examples throughout the co-labeling iterations, but it also employs an adaptive strategy in order to decide whether to trigger the editing operation according to different cases. The adaptive strategy combines five pre-conditional theorems, all of which ensure an iterative reduction of classification error and an increase in the scale of new training sets under PAC learning theory. Experiments on UCI datasets and an application to small pulmonary nodules detection using chest CT images show that ADE-Co-Forest can more effectively enhance the performance of a learned hypothesis than Co-Forest and DE-Co-Forest (Co-Forest with Data Editing but without adaptive strategy).
DOI: 10.4135/9781452244723.n136
发表时间: 1996
期刊: --
影响因子: --
作者:
P. Adriaans;Dolf Zantinge
通讯作者: P. Adriaans;Dolf Zantinge
DOI: 10.1109/tkde.2005.186
发表时间: 2005-11
影响因子: 8.9
作者:
Zhi-Hua Zhou;Ming Li
通讯作者: Zhi-Hua Zhou;Ming Li
DOI: 10.1016/s0167-8655(02)00225-8
发表时间: 2003-04
期刊: Pattern Recognit. Lett.
影响因子: --
作者:
J. S. Sánchez;R. Barandela;A. I. Marqués;R. Alejo;J. Badenas
通讯作者: J. S. Sánchez;R. Barandela;A. I. Marqués;R. Alejo;J. Badenas
DOI: 10.13053/cys-7-2-1056
发表时间: 2003-12
期刊: Computación y Sistemas
影响因子: --
作者:
Carmen Martínez;O. Fuentes
通讯作者: Carmen Martínez;O. Fuentes
DOI: --
发表时间: 1996
期刊: --
影响因子: --
作者:
C. Merz
通讯作者: C. Merz