Improving the classification performance of biological imbalanced datasets by swarm optimization algorithms

Improving the classification performance of biological imbalanced datasets by swarm optimization algorithms
复制标题

DOI:
10.1007/s11227-015-1541-6
复制
发表时间:
2016-10-01
影响因子:
3.3
通讯作者:
Fiaidhi, Jinan
Fiaidhi, Jinan
中科院分区:
计算机科学4区
文献类型:
--
作者:
Li, Jinyan;Fong, Simon;Fiaidhi, Jinan

文献摘要

被引文献

相似文献

分类是一种流行的有监督机器学习方法,在计算生物学中有许多应用,其中数据样本在数据挖掘的帮助下自动分类到预定义的标签。通常训练样本包含非常少的感兴趣的实例(例如,医学异常、人群中的罕见疾病和不寻常的综合征等),但也有很多正常的例子。目标标签之间的数据分布的这种不平衡比率阻碍了分类算法的有效性,因为诱导模型没有用足够数量的感兴趣标签的实例进行训练,而是被普通训练记录淹没。传统的补救措施试图重新平衡目标类的数据分布,通过人为地增加感兴趣的实例,减少大多数常见实例或两者的组合。虽然基本概念是有效的,但对于如何在制作稀有样本和减少规范之间取得平衡,以最大限度地提高分类准确性,没有明确的指导方针。在本文中,使用不同的群体策略(蝙蝠启发算法和PSO)的优化模型提出了自适应平衡类分布的增加/减少,根据生物数据集的属性。该优化被扩展以同时实现尽可能高的精度和Kappa统计。优化模型在五个不平衡的医学数据集上进行了测试,这些数据集来自肺手术日志和生物测定数据的虚拟筛选。计算机仿真结果表明,该优化模型在医学数据分类中的性能优于其他类平衡方法。
Classification which is a popular supervised machine learning method has many applications in computational biology, where data samples are automatically categorized into predefined labels with the aid of data mining. Often the training samples contain very few instances of interest (e.g., medical anomalies, rare disease in a population, and unusual syndromes, etc.), but many normal instances. Such imbalanced ratio of data distributions among the target labels hampers the efficacy of classification algorithms, because the induced model has not been trained with sufficient amount of instances of the interesting label(s), but overwhelmed with ordinary training records. Traditional remedies attempt to rebalance the data distributions of the target classes, by inflating the interesting instances artificially, reducing the majority of the common instances or a combination of both. Though the fundamental concept is effective, there is no clear guideline on how to strike a balance between fabricating the rare samples and reducing the norms, with the purpose of maximizing the classification accuracy. In this paper, an optimization model using different swarm strategies (Bat-inspired algorithm and PSO) is proposed for adaptively balancing the increase/decrease of the class distribution, depending on the properties of the biological datasets. The optimization is extended for achieving the highest possible accuracy and Kappa statistics at the same time as well. The optimization model is tested on five imbalanced medical datasets, which are sourced from lung surgery logs and virtual screening of bioassay data. Computer simulation results show that the proposed optimization model outperforms other class balancing methods in medical data classification.