Protein fold similarity estimated by a probabilistic approach based on Cα-Cα distance comparison

Protein fold similarity estimated by a probabilistic approach based on Cα-Cα distance comparison
复制标题

DOI:
10.1006/jmbi.2001.5250
复制
发表时间:
2002-01-25
影响因子:
5.6
通讯作者:
Pongor, S
Pongor, S
中科院分区:
生物学2区
文献类型:
--
作者:
Carugo, O;Pongor, S

文献摘要

被引文献

相似文献

由3到30个氨基酸残基分隔的残基之间的C-α-C-α距离的分布是蛋白质折叠的高度特征,这使得通过直接比较距离直方图来识别它们是可能的。通过列联表分析进行比较,得到相同的概率(骄傲分数),值在0到1之间。对于密切相关的结构,骄傲与C-α原子之间的均方根距离高度相关,但它提供了正确的分类,即使对于结构比对没有意义的无关结构也是如此。例如,对Cath折叠结构数据库的分析表明,根据获得的最高骄傲分数,98.8%的折叠属于正确的Cath同源超家族类别。结构比对和二级结构分配不是计算Pride所必需的,它的速度足够快,可以扫描大型数据库。(C)2002年学术出版社。
The distribution of the C-alpha-C-alpha distances between residues separated by three to 30 amino acid residues is highly characteristic of protein folds and makes it possible to identify them from a straightforward comparison of the distance histograms. The comparison is carried out by contingency table analysis and yields a probability of identity (PRIDE score), with values between zero and 1. For closely related structures, PRIDE is highly correlated with the root-mean-square distance between C-alpha atoms, but it provides a correct classification even for unrelated structures for which a structural alignment is not meaningful. For example, an analysis of the CATH database of fold structures showed that 98.8 % of the folds fall into the correct CATH homologous superfamily category, based on the highest PRIDE score obtained. Structural alignment and secondary-structure assignment are not necessary for the calculation of PRIDE, which is fast enough to allow the scanning of large databases. (C) 2002 Academic Press.