KNN-based dynamic query-driven sample rescaling strategy for class imbalance learning

KNN-based dynamic query-driven sample rescaling strategy for class imbalance learning
复制标题

基于KNN的动态查询驱动的类不平衡学习样本重缩放策略

DOI:
10.1016/j.neucom.2016.01.043
复制
发表时间:
2016-05-26
期刊:
影响因子:
6
通讯作者:
Yu, Dong-Jun
Yu, Dong-Jun
中科院分区:
计算机科学2区
文献类型:
--
作者:
Hu, Jun;Li, Yang;Yu, Dong-Jun

文献摘要

被引文献

相似文献

在生物信息学预测问题中,大多数样本的数量显著大于少数样本的数量,类不平衡现象是普遍存在的。缓解类不平衡的严重程度已被证明是一种有前途的途径,用于增强基于统计机器学习的预测器在不平衡学习场景下的预测性能。在这项研究中,我们提出了一种新的动态查询驱动的样本重新缩放(DQD-SR)的策略来解决类不平衡。与传统的样本缩放技术,这往往会产生一个固定的平衡数据集,建议DQD-SR动态生成一个查询驱动的平衡数据集的基础上KNN算法。在传统的样本重缩放(T-SR)衍生的平衡数据集上训练的预测模型将部分学习隐藏在原始数据集中的全局知识,而在DQD-SR上训练的预测模型将反映查询样本及其相关邻居之间的查询特定局部知识。因此,我们开发了一种集成方案,将基于T-SR的模型和基于DQD-SR的模型相结合,以进一步提高整体预测性能。为了证明所提出的方法的有效性,我们进行了严格的交叉验证和独立验证测试的基准数据集有关蛋白质核苷酸结合残基预测,这是一个典型的不平衡学习问题,在生物信息学。计算机实验结果表明,该方法具有较高的预测性能,优于现有的基于序列的蛋白质核苷酸结合残基预测器。我们还实现了一个名为TargetNUCs的预测器,可在http://csbio.njust.edu.cn/bioinf/TargetNUCs上免费获得。(C)2016爱思唯尔B. V.保留所有权利。
The class imbalance phenomenon is pervasive in bioinformatics prediction problems in which the number of majority samples is significantly larger than that of minority samples. Relieving the severity of class imbalance has been demonstrated to be a promising route for enhancing the prediction performance of a statistical machine learning-based predictor under an imbalanced learning scenario. In this study, we propose a novel dynamic query-driven sample rescaling (DQD-SR) strategy for addressing class imbalance. Unlike the traditional sample rescaling technique, which often yields a fixed balanced dataset, the proposed DQD-SR dynamically generates a query-driven balanced dataset based on KNN algorithm. A prediction model trained on a traditional sample rescaling (T-SR)-derived balanced dataset will partially learn the global knowledge buried in the original dataset, whereas a prediction model trained on DQD-SR will reflect the query-specific local knowledge between a query sample and its correlated neighbors in the original dataset. Thus, we developed an ensemble scheme to integrate the T-SR-based model and the DQD-SR-based model to further improve the overall prediction performance. To demonstrate the efficacy of the proposed method, we performed stringent cross-validation and independent validation tests on benchmark datasets concerning protein-nucleotide binding residues prediction, which is a typical imbalanced learning problem in bioinformatics. Computer experimental results show that the proposed method achieves high prediction performance and outperforms existing sequence-based protein nucleotide binding residues predictors. We also implemented a predictor called TargetNUCs, which is freely available for academic use at http://csbio.njust.edu.cn/bioinf/TargetNUCs. (C) 2016 Elsevier B.V. All rights reserved.