Predicting absolute contact numbers of native protein structure from amino acid sequence

Predicting absolute contact numbers of native protein structure from amino acid sequence
复制标题

DOI:
10.1002/prot.20300
复制
发表时间:
2005-01-01
影响因子:
2.9
通讯作者:
Nishikawa, K
Nishikawa, K
中科院分区:
生物学4区
文献类型:
--
作者:
Kinjo, AR;Horimoto, K;Nishikawa, K

文献摘要

被引文献

相似文献

蛋白质结构中氨基酸残基的接触数由给定残基的C-β原子周围的C-β原子的数量定义,该数量类似于但不同于溶剂可及表面积。我们提出了一种从蛋白质的氨基酸序列预测蛋白质的接触数的方法。该方法是基于一个简单的线性回归计划,并预测的绝对值的接触人数。当使用单个序列进行参数估计和交叉验证时,本方法预测的接触数的平均相关系数为0.555。当使用多序列比对时,相关性增加到0.627,这是对先前方法的显著改进。在离散状态预测方面,2-,3-和10-状态预测的准确度分别为71.4%,54.1%和18.9%与残基类型相关的无偏阈值,76.3%,59.2%和21.8%与残基类型无关的无偏阈值。从预测的角度讨论了可及表面积和接触数的区别,以及接触数预测在三维结构预测中的应用。(C)2004 Wiley-Liss,Inc.
The contact number of an amino acid residue in a protein structure is defined by the number of C-beta atoms around the C-beta atom of the given residue, a quantity similar to, but different from, solvent accessible surface area. We present a method to predict the contact numbers of a protein from its amino acid sequence. The method is based on a simple linear regression scheme and predicts the absolute values of contact numbers. When single sequences are used for both parameter estimation and cross-validation, the present method predicts the contact numbers with a correlation coefficient of 0.555 on average. When multiple sequence alignments are used, the correlation increases to 0.627, which is a significant improvement over previous methods. In terms of discrete states prediction, the accuracies for 2-, 3-, and 10-state predictions are, respectively, 71.4%, 54.1%, and 18.9% with residue type-dependent unbiased thresholds, and 76.3%, 59.2%, and 21.8% with residue type-independent unbiased thresholds. The difference between accessible surface area and contact number from a prediction viewpoint and the application of contact number prediction to three-dimensional structure prediction are discussed. (C) 2004 Wiley-Liss, Inc.