Transmembrane topology and signal peptide prediction using dynamic bayesian networks.

Transmembrane topology and signal peptide prediction using dynamic bayesian networks.
复制标题

DOI:
10.1371/journal.pcbi.1000213
复制
发表时间:
2008-11
影响因子:
4.3
通讯作者:
Noble WS
Noble WS
中科院分区:
生物学2区
文献类型:
--
作者:
Reynolds SM;Käll L;Riffle ME;Bilmes JA;Noble WS

文献摘要

参考文献

被引文献

相似文献

隐马尔可夫模型已成功地应用于跨膜蛋白质拓扑结构预测和信号肽预测。在本文中,我们扩展这项工作,利用更强大的动态贝叶斯网络(DBN)类。我们的模型,Philius,灵感来自于以前发表的HMM,Phobius,并结合了信号肽子模型与跨膜子模型。我们介绍了一个两阶段的DBN解码器,结合了后验解码的能力与维特比风格解码的语法约束。Philius还提供了蛋白质类型、片段和拓扑结构置信度指标,以帮助解释预测。我们报告了13%的相对改善超过Phobius在全拓扑预测跨膜蛋白的准确性,和0.96的灵敏度和特异性检测信号肽。我们还表明,我们的置信度与所观察到的精度相关。此外,我们还对酵母资源中心(YRC)数据库中的所有630万种蛋白质进行了预测。这项大规模的研究提供了包括信号肽和/或一个或多个跨膜片段的蛋白质的相对数量的全貌,也为科学界提供了宝贵的资源。所有DBN都是使用图形模型工具包实现的。这里描述的模型的源代码可以在http://noble.gs.washington.edu/proj/philius上获得。Philius网络服务器可在http://www.yeastrc.org/philius上查阅,YRC数据库的预测可在http://www.yeastrc.org/pdr上查阅。跨膜蛋白控制信息和物质进出细胞的流动,并参与广泛的生物过程。它们的相互作用使它们成为药物靶点,据估计,最近推出的药物中有50%以上靶向膜蛋白。然而,通过实验确定跨膜蛋白的三维结构仍然是一项困难的任务,并且尽管给定生物体中多达四分之一的蛋白质是跨膜蛋白,但目前已知的三级结构中很少有跨膜蛋白。因此,用于预测跨膜蛋白的基本拓扑结构的计算方法引起了极大的兴趣,并且这些方法必须能够区分成熟的跨膜蛋白和在首次合成时含有N-末端跨膜信号肽的蛋白。在这项工作中,我们提出了Philius,一种新的计算方法,在同时检测信号肽和正确预测跨膜蛋白的拓扑结构方面优于以前的方法。Philius还为每个预测提供了一组置信度得分。一个Philius Web服务器向公众开放,酵母资源中心数据库中有超过600万种蛋白质的预先计算预测。
Hidden Markov models (HMMs) have been successfully applied to the tasks of transmembrane protein topology prediction and signal peptide prediction. In this paper we expand upon this work by making use of the more powerful class of dynamic Bayesian networks (DBNs). Our model, Philius, is inspired by a previously published HMM, Phobius, and combines a signal peptide submodel with a transmembrane submodel. We introduce a two-stage DBN decoder that combines the power of posterior decoding with the grammar constraints of Viterbi-style decoding. Philius also provides protein type, segment, and topology confidence metrics to aid in the interpretation of the predictions. We report a relative improvement of 13% over Phobius in full-topology prediction accuracy on transmembrane proteins, and a sensitivity and specificity of 0.96 in detecting signal peptides. We also show that our confidence metrics correlate well with the observed precision. In addition, we have made predictions on all 6.3 million proteins in the Yeast Resource Center (YRC) database. This large-scale study provides an overall picture of the relative numbers of proteins that include a signal-peptide and/or one or more transmembrane segments as well as a valuable resource for the scientific community. All DBNs are implemented using the Graphical Models Toolkit. Source code for the models described here is available at http://noble.gs.washington.edu/proj/philius. A Philius Web server is available at http://www.yeastrc.org/philius, and the predictions on the YRC database are available at http://www.yeastrc.org/pdr. Transmembrane proteins control the flow of information and substances into and out of the cell and are involved in a broad range of biological processes. Their interfacing role makes them rewarding drug targets, and it is estimated that more than 50% of recently launched drugs target membrane proteins. However, experimentally determining the three-dimensional structure of a transmembrane protein is still a difficult task, and few of the currently known tertiary structures are of transmembrane proteins despite the fact that as many as one quarter of the proteins in a given organism are transmembrane proteins. Computational methods for predicting the basic topology of a transmembrane protein are therefore of great interest, and these methods must be able to distinguish between mature, membrane-spanning proteins and proteins that, when first synthesized, contain an N-terminal membrane-spanning signal peptide. In this work, we present Philius, a new computational approach that outperforms previous methods in simultaneously detecting signal peptides and correctly predicting the topology of transmembrane proteins. Philius also supplies a set of confidence scores with each prediction. A Philius Web server is available to the public as well as precomputed predictions for over six million proteins in the Yeast Resource Center database.
DOI: 10.1186/1471-2105-6-s4-s12
发表时间: 2005-12-01
期刊: BMC bioinformatics
影响因子: 3
作者:
Fariselli P;Martelli PL;Casadio R
通讯作者: Casadio R
Pongo:用于全α跨膜蛋白的多个预测的Web服务器。
DOI: 10.1093/nar/gkl208
发表时间: 2006-07-01
影响因子: 14.9
作者:
Amico, Mauro;Finelli, Michele;Rossi, Ivan;Zauli, Andrea;Elofsson, Arne;Viklund, Hakan;von Heijne, Gunnar;Jones, David;Krogh, Anders;Fariselli, Piero;Martelli, Pier Luigi;Casadio, Rita
通讯作者: Casadio, Rita
DOI: 10.1016/s0022-2836(03)00182-7
发表时间: 2003-03-28
影响因子: 5.6
作者:
Melén, K;Krogh, A;von Heijne, G
通讯作者: von Heijne, G
DOI: 10.1093/bioinformatics/18.4.617
发表时间: 2002-04-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Delorenzi, M;Speed, T
通讯作者: Speed, T
DOI: 10.1093/bioinformatics/17.7.646
发表时间: 2001-07-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Möller, S;Croning, MDR;Apweiler, R
通讯作者: Apweiler, R