Identifying MoRFs in Disordered Protein Using Enlarged Conserved Features

Identifying MoRFs in Disordered Protein Using Enlarged Conserved Features
复制标题

使用扩大的保守特征识别无序蛋白质中的 MoRF

DOI:
10.1145/3194480.3198908
复制
发表时间:
2018
期刊:
2018 6th International Conference on Bioinformatics and Computational Biology
影响因子:
--
通讯作者:
and K. Shimizu
and K. Shimizu
中科院分区:
--
文献类型:
--
作者:
Chun Fang;Yoshitaka Moriwaki;Daming Zhu;and K. Shimizu

文献摘要

相似文献

识别内在无序蛋白(IDPs)中的短结合区,即分子识别特征(morf),是理解内在无序蛋白功能、确定蛋白质结构和设计药物的关键步骤。由于IDPs的复杂性,从其氨基酸序列高度准确地预测morf仍然是极具挑战性的。在此,受信号处理技术的启发,我们提出了一种基于序列扩大保守性的morf预测新方法。在我们的方法中,仅使用由序列生成的修正位置特定评分矩阵(PSSM)作为输入特征,并采用支持向量机(SVM)构建预测模型。最后,对输出的预测分数进行平均策略处理,进一步提高准确率。当与其他基于单一模型的方法在相同的数据集上进行比较时,我们的结果在准确性方面与最先进的方法相比非常有竞争力。
Identifying the short binding regions, which are called molecular recognition features (MoRFs), within intrinsically disordered proteins (IDPs) is the key step for understanding the function of IDPs, for protein structure determination and for drug design. Due to the complexity of IDPs, highly accurate prediction of MoRFs from its amino acid sequence still remains extremely challenging. Here, inspired by the signal processing technology, we proposed a new method which is based on the enlarged conserved features of sequence for MoRFs prediction. In our approach, only the revised position-specific scoring matrix (PSSM) generated from the sequence was used as input feature, and the support vector machine (SVM) was adopted to build the prediction model. Finally, the output prediction scores were processed by an average strategy to further improve the accuracy. When compared with other single model-based methods on the same datasets, our results were very competitive in terms of accuracy with respect to the state-of-the-art methods.