A new supervised over-sampling algorithm with application to protein-nucleotide binding residue prediction.

A new supervised over-sampling algorithm with application to protein-nucleotide binding residue prediction.
复制标题

一种应用于蛋白质-核苷酸结合残基预测的新型监督过采样算法

DOI:
10.1371/journal.pone.0107676
复制
发表时间:
2014
期刊:
影响因子:
3.7
通讯作者:
Shen HB
Shen HB
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Hu J;He X;Yu DJ;Yang XB;Yang JY;Shen HB

文献摘要

参考文献

被引文献

相似文献

蛋白质-核苷酸相互作用普遍存在于多种生物过程中。准确地识别相互作用残基单独从蛋白质序列是有用的蛋白质功能注释和药物设计,特别是在后基因组时代,大量的蛋白质数据没有功能注释。蛋白质-核苷酸结合残基预测是一个典型的不平衡学习问题,其中结合残基的数量远少于非结合残基。减轻类不平衡的严重性已被证明是一种很有前途的手段,提高基于机器学习的预测器的类不平衡问题的预测性能。然而,很少有人注意到类不平衡对蛋白质-核苷酸结合残基预测的负面影响。在这项研究中,我们提出了一个新的监督过采样算法,合成额外的少数类样本,以解决类不平衡。在蛋白质-核苷酸相互作用数据集上的实验结果表明,所提出的监督过采样算法能够有效缓解类不平衡的严重程度,有助于提高预测性能.基于所提出的过采样算法,一个预测器,称为TargetSOS,实现蛋白质-核苷酸结合残基预测。交叉验证测试和独立验证测试证明了TargetSOS的有效性。本研究中使用的网络服务器和数据集可在http://www.csbio.sjtu.edu.cn/bioinf/TargetSOS/上免费获得。
Protein-nucleotide interactions are ubiquitous in a wide variety of biological processes. Accurately identifying interaction residues solely from protein sequences is useful for both protein function annotation and drug design, especially in the post-genomic era, as large volumes of protein data have not been functionally annotated. Protein-nucleotide binding residue prediction is a typical imbalanced learning problem, where binding residues are extremely fewer in number than non-binding residues. Alleviating the severity of class imbalance has been demonstrated to be a promising means of improving the prediction performance of a machine-learning-based predictor for class imbalance problems. However, little attention has been paid to the negative impact of class imbalance on protein-nucleotide binding residue prediction. In this study, we propose a new supervised over-sampling algorithm that synthesizes additional minority class samples to address class imbalance. The experimental results from protein-nucleotide interaction datasets demonstrate that the proposed supervised over-sampling algorithm can relieve the severity of class imbalance and help to improve prediction performance. Based on the proposed over-sampling algorithm, a predictor, called TargetSOS, is implemented for protein-nucleotide binding residue prediction. Cross-validation tests and independent validation tests demonstrate the effectiveness of TargetSOS. The web-server and datasets used in this study are freely available at http://www.csbio.sjtu.edu.cn/bioinf/TargetSOS/.
DOI: 10.1186/1471-2091-12-20
发表时间: 2011-05-13
期刊: BMC biochemistry
影响因子: --
作者:
Firoz A;Malik A;Joplin KH;Ahmad Z;Jha V;Ahmad S
通讯作者: Ahmad S
DOI: 10.1186/1471-2105-11-301
发表时间: 2010-06-03
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Chauhan, Jagat S.;Mishra, Nitish K.;Raghava, Gajendra P. S.
通讯作者: Raghava, Gajendra P. S.
DOI: 10.1186/1477-5956-9-s1-s4
发表时间: 2011-10-14
期刊: Proteome science
影响因子: 2
作者:
Chen K;Mizianty MJ;Kurgan L
通讯作者: Kurgan L
DOI: 10.1186/1471-2105-10-434
发表时间: 2009-12-19
期刊: BMC bioinformatics
影响因子: 3
作者:
Chauhan JS;Mishra NK;Raghava GP
通讯作者: Raghava GP
DOI: 10.1007/3-540-48229-6_9
发表时间: 2001-01-01
期刊: ARTIFICIAL INTELLIGENCE IN MEDICINE, PROCEEDINGS
影响因子: --
作者:
Laurikkala, J
通讯作者: Laurikkala, J