Identification and application of the concepts important for accurate and reliable protein secondary structure prediction

Identification and application of the concepts important for accurate and reliable protein secondary structure prediction
复制标题

DOI:
10.1002/pro.5560051116
复制
发表时间:
1996-11-01
期刊:
影响因子:
8
通讯作者:
Sternberg, MJE
Sternberg, MJE
中科院分区:
生物学3区
文献类型:
--
作者:
King, RD;Sternberg, MJE

文献摘要

被引文献

相似文献

提出了一种根据多重比对同源序列预测蛋白质二级结构的方法,每个残基三态的总体准确度为 70.1%。有两个目标:通过识别一组对预测很重要的概念,然后使用线性统计来获得高精度;为了深入了解折叠过程,二级结构预测中的重要概念被确定为:残基构象倾向、序列边缘效应、疏水矩、同源序列中插入和删除的位置、保守矩、自相关、残基比率、二级结构反馈效应和过滤。边缘效应、守恒矩和自相关的显式使用对本文来说是新的,通过逐步添加信息和检查判别函数中的权重来分析预测中使用的概念的相对重要性,预测的简单而明确的结构允许轻松地重新实现该方法,预测的准确性是先验可预测的,这允许评估预测的效用:10%的链 预测被正确识别为平均准确度 >80%。现有的高精度预测方法是基于复杂非线性统计的“黑盒”预测器(例如,PHD 中的神经网络:Rost & Sander,1993a)。对于中短链(大于或等于 90 个残基且
A protein secondary structure prediction method from multiply aligned homologous sequences is presented with an overall per residue three-state accuracy of 70.1%. There are two aims: to obtain high accuracy by identification of a set of concepts important for prediction followed by use of linear statistics; and to provide insight into the folding process, The important concepts in secondary structure prediction are identified as: residue conformational propensities, sequence edge effects, moments of hydrophobicity, position of insertions and deletions in aligned homologous sequence, moments of conservation, auto-correlation, residue ratios, secondary structure feedback effects, and filtering. Explicit use of edge effects, moments of conservation, and auto-correlation are new to this paper, The relative importance of the concepts used in prediction was analyzed by stepwise addition of information and examination of weights in the discrimination function, The simple and explicit structure of the prediction allows the method to be reimplemented easily, The accuracy of a prediction is predictable a priori, This permits evaluation of the utility of the prediction: 10% of the chains predicted were identified correctly as having a mean accuracy of >80%. Existing high-accuracy prediction methods are ''black-box'' predictors based on complex nonlinear statistics (e.g., neural networks in PHD: Rost & Sander, 1993a). For medium- to short-length chains (greater than or equal to 90 residues and