Adaptive swarm cluster-based dynamic multi-objective synthetic minority oversampling technique algorithm for tackling binary imbalanced datasets in biomedical data classification.

Adaptive swarm cluster-based dynamic multi-objective synthetic minority oversampling technique algorithm for tackling binary imbalanced datasets in biomedical data classification.
复制标题

DOI:
10.1186/s13040-016-0117-1
复制
发表时间:
2016
期刊:
影响因子:
4.5
通讯作者:
Wong KKL
Wong KKL
中科院分区:
生物学3区
文献类型:
--
作者:
Li J;Fong S;Sung Y;Cho K;Wong R;Wong KKL

文献摘要

参考文献

被引文献

相似文献

不平衡数据集被定义为在感兴趣和不感兴趣的类中具有不平衡数据比例的训练数据集。通常在生物医学应用中,刺激类的样本在人群中是罕见的,例如医学异常,阳性临床测试和特定疾病。虽然原始数据集中的目标样本数量很少,但由于少数类的训练不足,在这种训练数据上引入分类模型会导致预测性能不佳。在本文中,我们使用一种新的类平衡方法命名为自适应群聚类的动态多目标合成少数过采样技术(ASCB_DmSMOTE)来解决这个不平衡的数据集问题,这是常见的生物医学应用。该方法结合欠采样和过采样到一个群优化算法。它自适应地选择合适的参数的再平衡算法,以找到最佳的解决方案。与其他版本的SMOTE算法相比,ASCB_DmSMOTE算法具有显著的改进,包括更高的准确性和可信度。我们提出的方法巧妙地结合了两个再平衡技术在一起。该算法在细节上合理地重新分配多数类,并动态优化SMOTE的两个参数,为每个聚类的子不平衡数据集合成合理规模的少数类。所提出的方法最终克服了其他传统的方法,并取得了更高的可信度,甚至更高的分类模型的准确性。
An imbalanced dataset is defined as a training dataset that has imbalanced proportions of data in both interesting and uninteresting classes. Often in biomedical applications, samples from the stimulating class are rare in a population, such as medical anomalies, positive clinical tests, and particular diseases. Although the target samples in the primitive dataset are small in number, the induction of a classification model over such training data leads to poor prediction performance due to insufficient training from the minority class. In this paper, we use a novel class-balancing method named adaptive swarm cluster-based dynamic multi-objective synthetic minority oversampling technique (ASCB_DmSMOTE) to solve this imbalanced dataset problem, which is common in biomedical applications. The proposed method combines under-sampling and over-sampling into a swarm optimisation algorithm. It adaptively selects suitable parameters for the rebalancing algorithm to find the best solution. Compared with the other versions of the SMOTE algorithm, significant improvements, which include higher accuracy and credibility, are observed with ASCB_DmSMOTE. Our proposed method tactfully combines two rebalancing techniques together. It reasonably re-allocates the majority class in the details and dynamically optimises the two parameters of SMOTE to synthesise a reasonable scale of minority class for each clustered sub-imbalanced dataset. The proposed methods ultimately overcome other conventional methods and attains higher credibility with even greater accuracy of the classification model.
DOI: 10.1166/jmihi.2016.1807
发表时间: 2016-08-01
影响因子: --
作者:
Li, Jinyan;Fong, Simon;Tan, Zhen
通讯作者: Tan, Zhen
DOI: 10.1007/s11227-015-1541-6
发表时间: 2016-10-01
影响因子: 3.3
作者:
Li, Jinyan;Fong, Simon;Fiaidhi, Jinan
通讯作者: Fiaidhi, Jinan
DOI: 10.1016/j.patcog.2011.02.019
发表时间: 2011-08-01
影响因子: 8
作者:
Fernandez-Navarro, Francisco;Hervas-Martinez, Cesar;Antonio Gutierrez, Pedro
通讯作者: Antonio Gutierrez, Pedro
DOI: 10.1109/tsmcb.2008.2002909
发表时间: 2009-02-01
影响因子: --
作者:
Tang, Yuchun;Zhang, Yan-Qing;Krasser, Sven
通讯作者: Krasser, Sven
DOI: 10.2307/2529310
发表时间: 1977-01-01
期刊: BIOMETRICS
影响因子: 1.9
作者:
LANDIS, JR;KOCH, GG
通讯作者: KOCH, GG