A hybrid sampling algorithm combining M-SMOTE and ENN based on Random forest for medical imbalanced data

A hybrid sampling algorithm combining M-SMOTE and ENN based on Random forest for medical imbalanced data
复制标题

基于随机森林的M-SMOTE与ENN相结合的医疗不平衡数据混合采样算法

DOI:
10.1016/j.jbi.2020.103465
复制
发表时间:
2020-07-01
影响因子:
4.5
通讯作者:
Kou, Yue
Kou, Yue
中科院分区:
医学3区
文献类型:
--
作者:
Xu, Zhaozhao;Shen, Derong;Kou, Yue

文献摘要

被引文献

相似文献

数据不平衡的分类问题在医学诊断中经常存在。传统分类算法通常假定每个类别中的样本数量相近,且在训练过程中它们的误分类代价相等。然而,患者样本的误分类代价高于健康人样本。因此,如何在不影响健康个体分类的情况下提高对患者的识别是一个亟待解决的问题。为了解决医学诊断中数据不平衡分类的问题,我们提出了一种名为RFMSE的混合采样算法,它将基于随机森林(RF)的面向误分类的合成少数类过采样技术(M - SMOTE)和编辑最近邻(ENN)相结合。该算法主要由三部分组成。首先,使用M - SMOTE增加少数类样本的数量,同时M - SMOTE的过采样率是RF的误分类率。然后,使用ENN从多数类样本中去除噪声样本。最后,使用RF对混合采样后的样本进行分类预测,并根据分类指标(即马修斯相关系数(MCC))的变化确定迭代的停止准则。当MCC的值持续下降时,迭代过程将停止。在十个UCI数据集上进行的大量实验表明,RFMSE能够有效解决数据不平衡分类的问题。与传统算法相比,我们的方法能够更有效地提高F值和MCC。
The problem of imbalanced data classification often exists in medical diagnosis. Traditional classification algorithms usually assume that the number of samples in each class is similar and their misclassification cost during training is equal. However, the misclassification cost of patient samples is higher than that of healthy person samples. Therefore, how to increase the identification of patients without affecting the classification of healthy individuals is an urgent problem. In order to solve the problem of imbalanced data classification in medical diagnosis, we propose a hybrid sampling algorithm called RFMSE, which combines the Misclassification-oriented Synthetic minority over-sampling technique (M-SMOTE) and Edited nearset neighbor (ENN) based on Random forest (RF). The algorithm is mainly composed of three parts. First, M-SMOTE is used to increase the number of samples in the minority class, while the over-sampling rate of M-SMOTE is the misclassification rate of RF. Then, ENN is used to remove the noise ones from the majority samples. Finally, RF is used to perform classification prediction for the samples after hybrid sampling, and the stopping criterion for iterations is determined according to the changes of the classification index (i.e. Matthews Correlation Coefficient (MCC)). When the value of MCC continuously drops, the process of iterations will be stopped. Extensive experiments conducted on ten UCI datasets demonstrate that RFMSE can effectively solve the problem of imbalanced data classification. Compared with traditional algorithms, our method can improve F-value and MCC more effectively.