Sequence based prediction of DNA-binding proteins based on hybrid feature selection using random forest and Gaussian naïve Bayes.

Sequence based prediction of DNA-binding proteins based on hybrid feature selection using random forest and Gaussian naïve Bayes.
复制标题

DOI:
10.1371/journal.pone.0086703
复制
发表时间:
2014
期刊:
影响因子:
3.7
通讯作者:
Zhang H
Zhang H
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Lou W;Wang X;Chen F;Chen Y;Jiang B;Zhang H

文献摘要

参考文献

被引文献

相似文献

由于DNA结合蛋白在基因调控中的重要作用,开发一种有效的方法来测定DNA结合蛋白,正变得非常需要,因为这将是非常宝贵的,以促进我们对蛋白质功能的理解。在这项研究中,我们提出了一种新的方法来预测的DNA结合蛋白质,通过执行的特征排名使用随机森林和包装器为基础的特征选择使用前向最佳优先搜索策略。这些特征包括来自一级序列、预测的二级结构、预测的相对溶剂可及性和位置特异性评分矩阵的信息。所提出的方法称为DBPPred,使用高斯朴素贝叶斯作为底层分类器,因为它优于其他五种分类器,包括决策树,逻辑回归,k最近邻,多项式核支持向量机和径向基函数支持向量机。因此,根据在训练基准数据集PDB 594上运行10次的五重交叉验证,所提出的DBPPred产生了0.791的最高平均准确度和0.583的平均MCC。随后,通过在整个PDB 594数据集上训练的所提出的模型和其他五种现有方法(包括iDNA-Prot、DNA-Prot、DNAbinder、DNABIND和DBD-Threader)对独立数据集PDB 186进行盲测,结果所提出的DBPPred产生了最高的准确度0.769,MCC为0.538,AUC为0.790。与相关预测方法相比,所提出的DBPPred对完全大型非DNA结合蛋白数据集和两个RNA结合蛋白数据集进行的独立测试也显示出改进或相当的质量。此外,我们观察到,大多数所提出的方法选择的功能之间的平均特征值的DNA结合和非DNA结合蛋白质的统计显着不同。所有的实验结果表明,提出的DBPPred可以是一个替代的前景预测大规模测定的DNA结合蛋白。
Developing an efficient method for determination of the DNA-binding proteins, due to their vital roles in gene regulation, is becoming highly desired since it would be invaluable to advance our understanding of protein functions. In this study, we proposed a new method for the prediction of the DNA-binding proteins, by performing the feature rank using random forest and the wrapper-based feature selection using forward best-first search strategy. The features comprise information from primary sequence, predicted secondary structure, predicted relative solvent accessibility, and position specific scoring matrix. The proposed method, called DBPPred, used Gaussian naïve Bayes as the underlying classifier since it outperformed five other classifiers, including decision tree, logistic regression, k-nearest neighbor, support vector machine with polynomial kernel, and support vector machine with radial basis function. As a result, the proposed DBPPred yields the highest average accuracy of 0.791 and average MCC of 0.583 according to the five-fold cross validation with ten runs on the training benchmark dataset PDB594. Subsequently, blind tests on the independent dataset PDB186 by the proposed model trained on the entire PDB594 dataset and by other five existing methods (including iDNA-Prot, DNA-Prot, DNAbinder, DNABIND and DBD-Threader) were performed, resulting in that the proposed DBPPred yielded the highest accuracy of 0.769, MCC of 0.538, and AUC of 0.790. The independent tests performed by the proposed DBPPred on completely a large non-DNA binding protein dataset and two RNA binding protein datasets also showed improved or comparable quality when compared with the relevant prediction methods. Moreover, we observed that majority of the selected features by the proposed method are statistically significantly different between the mean feature values of the DNA-binding and the non DNA-binding proteins. All of the experimental results indicate that the proposed DBPPred can be an alternative perspective predictor for large-scale determination of DNA-binding proteins.
DOI: 10.1093/nar/gki949
发表时间: 2005
影响因子: 14.9
作者:
Bhardwaj N;Langlois RE;Zhao G;Lu H
通讯作者: Lu H
DOI: 10.1002/jcc.21968
发表时间: 2012-01-30
影响因子: 3
作者:
Faraggi, Eshel;Zhang, Tuo;Yang, Yuedong;Kurgan, Lukasz;Zhou, Yaoqi
通讯作者: Zhou, Yaoqi
DOI: 10.1093/nar/gkn332
发表时间: 2008-07
影响因子: 14.9
作者:
Gao, Mu;Skolnick, Jeffrey
通讯作者: Skolnick, Jeffrey
DOI: 10.1093/nar/gks405
发表时间: 2012-08
影响因子: 14.9
作者:
Dey S;Pal A;Guharoy M;Sonavane S;Chakrabarti P
通讯作者: Chakrabarti P
DOI: 10.1016/j.febslet.2007.01.086
发表时间: 2007-03-06
期刊: FEBS LETTERS
影响因子: 3.5
作者:
Bhardwaj, Nitin;Lu, Hui
通讯作者: Lu, Hui