Imbalanced biomedical data classification using self-adaptive multilayer ELM combined with dynamic GAN

Imbalanced biomedical data classification using self-adaptive multilayer ELM combined with dynamic GAN
复制标题

DOI:
10.1186/s12938-018-0604-3
复制
发表时间:
2018-12-04
影响因子:
3.9
通讯作者:
Jiang, Zhengang
Jiang, Zhengang
中科院分区:
工程技术3区
文献类型:
--
作者:
Zhang, Liyuan;Yang, Huamin;Jiang, Zhengang

文献摘要

被引文献

相似文献

背景:不平衡数据分类是医学智能诊断中不可避免的问题。现实世界中的生物医学数据集通常是沿着的,具有有限的样本和高维特征。这严重影响了模型的分类性能,并导致对疾病诊断的错误指导。探索一种有效的分类方法,不平衡和有限的生物医学dataset是一个具有挑战性的task.Methods:在本文中,我们提出了一种新的多层极端学习机(ELM)分类模型结合动态生成对抗网络(GAN)来处理有限和不平衡的生物医学数据。首先,利用主成分分析去除不相关和冗余的特征。同时,提取出更有意义的病理特征。在此基础上,设计了动态GAN算法,生成具有真实感的少数类样本,从而平衡了类分布,有效地避免了过拟合。最后,提出了一种自适应多层ELM来对平衡数据集进行分类。通过定量地建立不平衡率变化与模型超参数之间的关系,确定了隐层数和节点数的解析表达式。结果:在4个真实生物医学数据集上进行了数值实验,验证了该方法的分类性能。该方法能够生成真实的少数类样本,并能自适应地选择最优的学习模型参数.通过与W-ELM、SMOTE-ELM和H-ELM方法的比较,定量实验结果表明,该方法在ROC、AUC、G-mean和F-measure等指标上均具有更好的分类性能和更高的计算效率。结论:该方法为有限样本和高维特征条件下的不平衡生物医学数据分类提供了一种有效的解决方案。该方法可为计算机辅助诊断提供理论依据。它具有在生物医学临床实践中应用的潜力。
Background: Imbalanced data classification is an inevitable problem in medical intelligent diagnosis. Most of real-world biomedical datasets are usually along with limited samples and high-dimensional feature. This seriously affects the classification performance of the model and causes erroneous guidance for the diagnosis of diseases. Exploring an effective classification method for imbalanced and limited biomedical dataset is a challenging task.Methods: In this paper, we propose a novel multilayer extreme learning machine(ELM) classification model combined with dynamic generative adversarial net (GAN) to tackle limited and imbalanced biomedical data. Firstly, principal component analysis is utilized to remove irrelevant and redundant features. Meanwhile, more meaningful pathological features are extracted. After that, dynamic GAN is designed to generate the realistic-looking minority class samples, thereby balancing the class distribution and avoiding overfitting effectively. Finally, a self-adaptive multilayer ELM is proposed to classify the balanced dataset. The analytic expression for the numbers of hidden layer and node is determined by quantitatively establishing the relationship between the change of imbalance ratio and the hyper-parameters of the model. Reducing interactive parameters adjustment makes the classification model more robust.Results: To evaluate the classification performanceof the proposed method, numerical experiments are conducted on four real-world biomedical datasets. Theproposed method can generate authentic minority class samples and self-adaptivelyselect the optimal parameters of learning model. By comparing with W-ELM, SMOTE-ELM, and H-ELM methods, the quantitative experimental results demonstrate that our method can achieve better classification performance and higher computational efficiency in terms of ROC, AUC, G-mean, and F-measure metrics.Conclusions: Our study provides an effective solution for imbalanced biomedical data classification under the condition of limited samples and high-dimensional feature. The proposed method could offer a theoretical basis for computer-aided diagnosis. Ithas the potential to be applied in biomedical clinical practice.