Solving the Under-Fitting Problem for Decision Tree Algorithms by Incremental Swarm Optimization in Rare-Event Healthcare Classification

Solving the Under-Fitting Problem for Decision Tree Algorithms by Incremental Swarm Optimization in Rare-Event Healthcare Classification
复制标题

DOI:
10.1166/jmihi.2016.1807
复制
发表时间:
2016-08-01
影响因子:
--
通讯作者:
Tan, Zhen
Tan, Zhen
中科院分区:
医学4区
文献类型:
--
作者:
Li, Jinyan;Fong, Simon;Tan, Zhen

文献摘要

被引文献

相似文献

众所周知,医疗保健数据在目标类的数据分布中是不平衡的,其中感兴趣的样本比普通样本少得多。在医疗数据分类方面,决策树归纳中的监督训练不足很容易发生,导致分类/预测准确性差。提出了一种基于群体平衡算法(SBA)的数据再平衡方法--合成少数过采样技术(SMOTE),该方法用于校正欠拟合问题。虽然SBA工作得很好,但它的缺点是要求所有数据必须在最初可用。在本文中,一个替代的方法,扩展了SBA,即增量群体平衡算法(ISBA)的决策树的影响进行了研究。ISBA通过优化SMOTE和动态训练决策树,以比SBA更快的速度获得更高的分类准确率。在我们的设计中,两个群体算法,粒子群优化和蝙蝠启发算法,用于耦合两种不同类型的决策树分类器,决策树(DT)和Hoeffding树(HT)。前者代表了传统的批式决策树模型,后者是典型的增量式决策树模型。通过两组不平衡的医疗数据进行实验,目的是比较和对比ISBA对DT和HT的有效性。
Healthcare data are well-known to be imbalanced in the data distribution of target classes where the samples of interest are much fewer than the ordinary samples. When it comes to healthcare data classification, insufficient supervised training in decision tree induction is prone to happen, leading to poor classification/prediction accuracy. Swarm Balancing Algorithm (SBA) was proposed to optimize the parameter values of a popular data-rebalancing method called Synthetic Minority Over-sampling Technique (SMOTE) for rectifying the under-fitting problems. Though it works well, the drawback of SBA is the requirement that all the data must be initially available. In this paper, an alternative approach which extends from SBA, namely, Incremental Swarm Balancing Algorithm (ISBA) is investigated on the impacts of decision trees. ISBA obtains higher classification accuracy at faster speed than SBA by optimizing SMOTE and training a decision tree on the fly. In our design, two swarm algorithms, particle swarm optimization and bat-inspired algorithm, are used to couple with two different types of decision tree classifiers, Decision Tree (DT) and Hoeffding Tree (HT). The former represents the traditional batch-type decision tree model, and the latter is typical incremental decision tree model. Experimentation over two sets of imbalanced healthcare data is performed, with the aim of comparing and contrasting the efficacy of ISBA for DT and HT.