DP-BINDER: machine learning model for prediction of DNA-binding proteins by fusing evolutionary and physicochemical information

DP-BINDER: machine learning model for prediction of DNA-binding proteins by fusing evolutionary and physicochemical information
复制标题

DOI:
10.1007/s10822-019-00207-x
复制
发表时间:
2019-07-01
影响因子:
3.5
通讯作者:
Akbar, Shahid
Akbar, Shahid
中科院分区:
生物学3区
文献类型:
--
作者:
Ali, Farman;Ahmed, Saeed;Akbar, Shahid

文献摘要

被引文献

相似文献

DNA 结合蛋白 (DBP) 参与各种生物过程,包括 DNA 复制、重组和修复。在人类基因组中,大约 6-7% 的这些蛋白质被用于基因编码。 DBP 将 DNA 塑造成一种称为染色质的紧凑结构,而其中一些蛋白质则调节染色体包装和转录过程。在制药行业,DBP 被用作抗生素、类固醇和抗癌药物的关键成分。这些蛋白质还参与 DNA 的生物物理、生物学和生化研究。由于 DBP 在多种生物活性中发挥着至关重要的作用,DBP 的鉴定是蛋白质科学中的一个热点问题。人们提出了一系列的实验和计算方法,但有些方法没有达到预期的结果,有些方法的准确性和真实性不够。尽管如此,仍然非常希望提供更智能的计算预测器。在这项工作中,我们介绍了一种基于物理化学和进化信息的创新计算方法,即 DP-BINDER。我们使用训练和独立数据集,通过归一化莫罗-布罗托自相关(NMBAC)从初级蛋白质序列的理化性质中捕获局部高度决定性的特征,并通过位置特定评分矩阵-转移概率组合(PSSM-TPC)和伪位置特定评分矩阵(PsePSSM)捕获进化信息。通过支持向量机-递归特征消除和相关偏差减少(SVM-RFE + CBR)从融合特征中选择最佳特征,并将其输入随机森林(RF)和支持向量机(SVM)。我们的方法在训练数据集上通过折刀和十倍交叉验证分别获得了 92.46% 和 89.58% 的准确率,而在独立数据集上预测 DBP 的准确率则达到了 81.17%。这些结果表明我们的方法取得了文献中最高的成功率。 DP-BINDER 相对于现有方法的优越性有几个原因,包括通过有效的特征描述符抽象局部主导特征、利用适当的特征选择算法和有效的分类器。
DNA-binding proteins (DBPs) participate in various biological processes including DNA replication, recombination, and repair. In the human genome, about 6-7% of these proteins are utilized for genes encoding. DBPs shape the DNA into a compact structure known chromatin while some of these proteins regulate the chromosome packaging and transcription process. In the pharmaceutical industry, DBPs are used as a key component of antibiotics, steroids, and cancer drugs. These proteins also involve in biophysical, biological, and biochemical studies of DNA. Due to the crucial role in various biological activities, identification of DBPs is a hot issue in protein science. A series of experimental and computational methods have been proposed, however, some methods didn't achieve the desired results while some are inadequate in its accuracy and authenticity. Still, it is highly desired to present more intelligent computational predictors. In this work, we introduce an innovative computational method namely DP-BINDER based on physicochemical and evolutionary information. We captured local highly decisive features from physicochemical properties of primary protein sequences via normalized Moreau-Broto autocorrelation (NMBAC) and evolutionary information by position specific scoring matrix-transition probability composition (PSSM-TPC) and pseudo-position specific scoring matrix (PsePSSM) using training and independent datasets. The optimal features were selected by the support vector machine-recursive feature elimination and correlation bias reduction (SVM-RFE + CBR) from fused features and were fed into random forest (RF) and support vector machine (SVM). Our method attained 92.46% and 89.58% accuracy with jackknife and ten-fold cross-validation, respectively on the training dataset, while 81.17% accuracy on the independent dataset for prediction of DBPs. These results demonstrate that our method attained the highest success rate in the literature. The superiority of DP-BINDER over existing approaches due to several reasons including abstraction of local dominant features via effective feature descriptors, utilization of appropriate feature selection algorithms and effective classifier.