DNA binding protein identification by combining pseudo amino acid composition and profile-based protein representation.

DNA binding protein identification by combining pseudo amino acid composition and profile-based protein representation.
复制标题

DOI:
10.1038/srep15479
复制
发表时间:
2015-10-20
期刊:
影响因子:
4.6
通讯作者:
Wang X
Wang X
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Liu B;Wang S;Wang X

文献摘要

被引文献

相似文献

DNA结合蛋白在大多数细胞过程中起重要作用。因此,有必要开发一种仅基于蛋白质序列信息的DNA结合蛋白的有效预测器。构建一个有用的预测器的瓶颈是找到合适的特征捕捉的DNA结合蛋白的特性。我们将PseAAC应用于DNA结合蛋白的识别,并通过使用基于profile-based蛋白质表示的方法结合进化信息对PseAAC进行了进一步的改进。最后,结合支持向量机(SVM),提出了一种预测器iDNAPro-PseAAC。在更新的基准数据集上的实验结果表明,iDNAPro-PseAAC优于一些最先进的方法,并且它可以在独立数据集上实现稳定的性能。通过使用集成学习方法在训练过程中引入更多的阴性样本(非DNA结合蛋白),iDNAPro-PseAAC的性能得到进一步提高。iDNAPro-PseAAC的网络服务器可在http://bioinformatics.hitsz.edu.cn/iDNAPro-PseAAC/上获得。
DNA-binding proteins play an important role in most cellular processes. Therefore, it is necessary to develop an efficient predictor for identifying DNA-binding proteins only based on the sequence information of proteins. The bottleneck for constructing a useful predictor is to find suitable features capturing the characteristics of DNA binding proteins. We applied PseAAC to DNA binding protein identification, and PseAAC was further improved by incorporating the evolutionary information by using profile-based protein representation. Finally, Combined with Support Vector Machines (SVMs), a predictor called iDNAPro-PseAAC was proposed. Experimental results on an updated benchmark dataset showed that iDNAPro-PseAAC outperformed some state-of-the-art approaches, and it can achieve stable performance on an independent dataset. By using an ensemble learning approach to incorporate more negative samples (non-DNA binding proteins) in the training process, the performance of iDNAPro-PseAAC was further improved. The web server of iDNAPro-PseAAC is available at http://bioinformatics.hitsz.edu.cn/iDNAPro-PseAAC/.