Accurate prediction of protein secondary structure and solvent accessibility by consensus combiners of sequence and structure information.

Accurate prediction of protein secondary structure and solvent accessibility by consensus combiners of sequence and structure information.
复制标题

DOI:
10.1186/1471-2105-8-201
复制
发表时间:
2007-06-14
期刊:
影响因子:
3
通讯作者:
Vullo A
Vullo A
中科院分区:
生物学4区
文献类型:
--
作者:
Pollastri G;Martin AJ;Mooney C;Vullo A

文献摘要

参考文献

被引文献

相似文献

蛋白质的结构特性(例如二级结构和溶剂可及性)有助于三维结构预测,不仅在从头算情况下,而且在已知结构的同源性信息可用时也是如此。即使同源性是可用的,结构特性也通常用于蛋白质分析,主要是因为同源性建模的通量低于二级结构预测。尽管如此,二级结构和溶剂可及性的预测几乎总是从头开始。在这里,我们开发了高通量机器学习系统,用于预测蛋白质二级结构和溶剂可及性,该系统利用与已知结构的蛋白质的同源性,在可用的情况下,以从PDB模板组中提取的简单结构频率曲线的形式。我们比较这些系统,他们的国家的最先进的从头算同行,并与一些基线中的二级结构和溶剂accessories直接从模板中提取。我们发现,模板的结构信息大大提高了二级结构和溶剂可及性预测质量,平均而言,系统显着丰富的模板中包含的信息。对于超过30%的序列相似性,二级结构预测质量约为90%,接近其理论最大值,2级溶剂可及性约为85%。增益相对于模板选择噪声是稳健的,并且对于边缘序列相似性和短比对是显著的,支持这些改进的预测可能证明在其中可获得明确同源性的情况之外是有益的这一主张。预测系统在地址公开提供。
Structural properties of proteins such as secondary structure and solvent accessibility contribute to three-dimensional structure prediction, not only in the ab initio case but also when homology information to known structures is available. Structural properties are also routinely used in protein analysis even when homology is available, largely because homology modelling is lower throughput than, say, secondary structure prediction. Nonetheless, predictors of secondary structure and solvent accessibility are virtually always ab initio. Here we develop high-throughput machine learning systems for the prediction of protein secondary structure and solvent accessibility that exploit homology to proteins of known structure, where available, in the form of simple structural frequency profiles extracted from sets of PDB templates. We compare these systems to their state-of-the-art ab initio counterparts, and with a number of baselines in which secondary structures and solvent accessibilities are extracted directly from the templates. We show that structural information from templates greatly improves secondary structure and solvent accessibility prediction quality, and that, on average, the systems significantly enrich the information contained in the templates. For sequence similarity exceeding 30%, secondary structure prediction quality is approximately 90%, close to its theoretical maximum, and 2-class solvent accessibility roughly 85%. Gains are robust with respect to template selection noise, and significant for marginal sequence similarity and for short alignments, supporting the claim that these improved predictions may prove beneficial beyond the case in which clear homology is available. The predictive system are publicly available at the address .
DOI: 10.1002/prot.340230412
发表时间: 1995-12-01
期刊: PROTEINS-STRUCTURE FUNCTION AND GENETICS
影响因子: --
作者:
Frishman, D;Argos, P
通讯作者: Argos, P
DOI: 10.1186/1471-2105-7-301
发表时间: 2006-06-14
期刊: BMC BIOINFORMATICS
影响因子: 3
作者:
Montgomerie, Scott;Sundararaj, Shan;Wishart, David S.
通讯作者: Wishart, David S.
DOI: 10.1002/prot.10556
发表时间: 2003-01-01
影响因子: 2.9
作者:
Moult, J;Fidelis, K;Hubbard, T
通讯作者: Hubbard, T
DOI: 10.1006/jmbi.1999.3091
发表时间: 1999-09-17
影响因子: 5.6
作者:
Jones, DT
通讯作者: Jones, DT
DOI: 10.1093/bioinformatics/bti1004
发表时间: 2005-06-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Cheng, JL;Baldi, P
通讯作者: Baldi, P