Computational Identification of MoRFs in Protein Sequences Using Hierarchical Application of Bayes Rule.

Computational Identification of MoRFs in Protein Sequences Using Hierarchical Application of Bayes Rule.
复制标题

使用贝叶斯规则的分层应用对蛋白质序列中MORF的计算鉴定。

DOI:
10.1371/journal.pone.0141603
复制
发表时间:
2015
期刊:
影响因子:
3.7
通讯作者:
Gsponer J
Gsponer J
中科院分区:
综合性期刊3区
文献类型:
--
作者:
Malhis N;Wong ET;Nassar R;Gsponer J

文献摘要

被引文献

相似文献

蛋白质的内在无序区域在各种生物过程的调节中起着重要作用。其调控功能的关键通常是通过称为分子识别特征(morf)的序列元件与球状蛋白结构域的结合。开发用于识别氨基酸序列中候选MoRF位置的计算工具是一项重要的任务,也是一个越来越受关注的领域。鉴于蛋白质序列中MoRF的相对稀疏性,可用的MoRF预测因子的准确性通常不足以用于实际使用,这留下了显著的需求和改进空间。在这项工作中,我们引入了MoRFCHiBi_Web,与目前的MoRF预测器相比,它预测蛋白质序列中MoRF的位置具有更高的准确性。三个不同的和很大程度上独立的属性得分计算与组成预测,然后结合产生最终的MoRF倾向得分。第一个分数反映了序列窗口包含morf的可能性,并基于氨基酸组成和序列相似性信息。它是由MoRFCHiBi使用多达40个残基大小的小窗口生成的。第二个分数用于识别蛋白质紊乱的长片段,由ESpritz和DisProt选项生成。最后,第三个分数反映了残数守恒,是由PSI-BLAST生成的PSSM文件组装而成的。对这些倾向得分进行处理,然后使用贝叶斯规则分层组合,以生成最终的MoRFCHiBi_Web预测。MoRFCHiBi_Web在三个数据集上进行了测试。结果表明,MoRFCHiBi_Web优于先前开发的预测器,在实际阈值下,对于相同的真阳性率,产生的假阳性率不到一半。这种精度水平与其相对较高的处理速度相结合,使MoRFCHiBi_Web成为MoRF预测的实用工具。http://morf.chibi.ubc.ca: 8080 / morf /。
Intrinsically disordered regions of proteins play an essential role in the regulation of various biological processes. Key to their regulatory function is often the binding to globular protein domains via sequence elements known as molecular recognition features (MoRFs). Development of computational tools for the identification of candidate MoRF locations in amino acid sequences is an important task and an area of growing interest. Given the relative sparseness of MoRFs in protein sequences, the accuracy of the available MoRF predictors is often inadequate for practical usage, which leaves a significant need and room for improvement. In this work, we introduce MoRFCHiBi_Web, which predicts MoRF locations in protein sequences with higher accuracy compared to current MoRF predictors. Three distinct and largely independent property scores are computed with component predictors and then combined to generate the final MoRF propensity scores. The first score reflects the likelihood of sequence windows to harbour MoRFs and is based on amino acid composition and sequence similarity information. It is generated by MoRFCHiBi using small windows of up to 40 residues in size. The second score identifies long stretches of protein disorder and is generated by ESpritz with the DisProt option. Lastly, the third score reflects residue conservation and is assembled from PSSM files generated by PSI-BLAST. These propensity scores are processed and then hierarchically combined using Bayes rule to generate the final MoRFCHiBi_Web predictions. MoRFCHiBi_Web was tested on three datasets. Results show that MoRFCHiBi_Web outperforms previously developed predictors by generating less than half the false positive rate for the same true positive rate at practical threshold values. This level of accuracy paired with its relatively high processing speed makes MoRFCHiBi_Web a practical tool for MoRF prediction. http://morf.chibi.ubc.ca:8080/morf/.