Prediction of distant residue contacts with the use of evolutionary information

Prediction of distant residue contacts with the use of evolutionary information
复制标题

DOI:
10.1002/prot.20370
复制
发表时间:
2005-03-01
影响因子:
2.9
通讯作者:
Kaznessis, Y
Kaznessis, Y
中科院分区:
生物学4区
文献类型:
--
作者:
Vicatos, S;Reddy, BVB;Kaznessis, Y

文献摘要

被引文献

相似文献

在这项工作中,我们提出了一种新的相关突变分析(CMA)的方法,是显着更准确的比以前报道的CMA方法。相关系数的计算是基于物理化学。性质的残留物(预测),而不是替代矩阵。这导致在蛋白质序列中远离但在其三维三级结构中接近的残基对的可靠预测。多重序列比对(MSA)包含一个序列的已知结构的127个家庭从PFAM数据库已被选定,使所有主要的蛋白质结构中描述的CATH分类数据库的代表。过滤所选家族中的蛋白质序列,使得只有那些进化上接近靶蛋白的蛋白质保留在MSA中。α-β类蛋白质的平均准确度为预测的近端对的26.8%,平均改善随机准确度(IOR)为6.41。平均准确率为20.6%的主要β类和14.4%的主要α类。发现最佳相关系数截止值(cc截止值)约为0.65。与疏水性相关的第一个预测因子提供了最可靠的结果。另外两个预测因子给出了很好的预测,可以与第一个预测因子结合使用。当选择更严格的cc截止值时,平均准确率显着增加(alpha beta类为38.76%),但代价是预测数量较少。使用溶剂可及面积估计过滤假阳性的预测是有前途的。(C)2005 Wiley-Liss,Inc.
In this work we present a novel correlated mutations analysis (CMA) method that is significantly more accurate than previously reported CMA methods. Calculation of correlation coefficients is based on physicochemical. properties of residues (predictors) and not on substitution matrices. This results in reliable prediction of pairs of residues that are distant in protein sequence but proximal in its three dimensional tertiary structure. Multiple sequence alignments (MSA) containing a sequence of known structure for 127 families from PFAM database have been selected so that all major protein architectures described in CATH classification database are represented. Protein sequences in the selected families were filtered so that only those evolutionarily close to the target protein remain in the MSA. The average accuracy obtained for the alpha beta class of proteins was 26.8% of predicted proximal pairs with average improvement over random accuracy (IOR) of 6.41. Average accuracy is 20.6% for the mainly beta class and 14.4% for the mainly alpha class. The optimum correlation coefficient cutoff (cc cutoff) was found to be around 0.65. The first predictor, which correlates to hydrophobicity, provides the most reliable results. The other two predictors give good predictions which can be used in conjunction to those of the first one. When stricter cc cutoff is chosen, the average accuracy increases significantly (38.76% for alpha beta class), but the trade off is a smaller number of predictions. The use of solvent accessible area estimations for filtering false positives out of the predictions is promising. (C) 2005 Wiley-Liss, Inc.