The PDB is a covering set of small protein structures

The PDB is a covering set of small protein structures
复制标题

DOI:
10.1016/j.jmb.2003.10.027
复制
发表时间:
2003-12-05
影响因子:
5.6
通讯作者:
Skolnick, J
Skolnick, J
中科院分区:
生物学2区
文献类型:
--
作者:
Kihara, D;Skolnick, J

文献摘要

被引文献

相似文献

所有代表性蛋白质的结构比较都已完成。采用与天然的相对均方根偏差(RMSD)使得能够根据Z分数评估不同长度的结构比对的统计显著性。得出两个结论:第一,具有天然折叠的蛋白质可以通过它们的Z分数来区分。其次,有点令人惊讶的是,所有长度达100个残基的小蛋白质与不同二级结构和折叠类别的其他蛋白质具有显着的结构比对;即其中24.0%的蛋白质被模板蛋白质覆盖60%,RMSD低于3.5埃,6.0%的蛋白质覆盖率为70%。如果我们仅比对具有不同二级结构类型的蛋白质的限制被去除,则在200个残基或更小的蛋白质的代表性基准组中,93%可以与单个模板结构(具有9.8%的平均序列同一性)比对,具有小于4埃的RMSD和79%的平均覆盖度。从这个意义上说,目前的蛋白质数据库(PDB)几乎是一个覆盖小蛋白质结构的集合。比对区域的长度(相对于整个蛋白质长度)在最高命中蛋白质之间没有差异,表明蛋白质结构空间是高度密集的。对于较大的蛋白质,非相关蛋白质可以覆盖结构的重要部分。此外,这些顶级命中蛋白质与靶蛋白的不同部分对齐,因此组合时几乎可以覆盖整个分子。覆盖靶蛋白所需的蛋白质数量非常少,例如,对于长达320个残基的蛋白质,前十个命中蛋白质可以给出低于3.5 A的RMSD的90%覆盖率。这些结果给出了一个新的观点的蛋白质结构空间的性质,并讨论了其对蛋白质结构预测的影响。(C)2003 Elsevier Ltd.保留所有权利。
Structure comparisons of all representative proteins have been done. Employing the relative root mean square deviation (RMSD) from native enables the assessment of the statistical significance of structure alignments of different lengths in terms of a Z-score. Two conclusions emerge: first, proteins with their native fold can be distinguished by their Z-score. Second and somewhat surprising, all small proteins up to 100 residues in length have significant structure alignments to other proteins in a different secondary structure and fold class; i.e. 24.0% of them have 60% coverage by a template protein with a RMSD below 3.5 Angstrom and 6.0% have 70% coverage. If the restriction that we align proteins only having different secondary structure types is removed, then in a representative benchmark set of proteins of 200 residues or smaller, 93% can be aligned to a single template structure (with average sequence identity of 9.8%), with a RMSD less than 4 Angstrom, and 79% average coverage. In this sense, the current Protein Data Bank (PDB) is almost a covering set of small protein structures. The length of the aligned region (relative to the whole protein length) does not differ among the top hit proteins, indicating that protein structure space is highly dense. For larger proteins, non-related proteins can cover a significant portion of the structure. Moreover, these top hit proteins are aligned to different parts of the target protein, so that almost the entire molecule can be covered when combined. The number of proteins required to cover a target protein is very small, e.g. the top ten hit proteins can give 90% coverage below a RMSD of 3.5 A for proteins up to 320 residues long. These results give a new view of the nature of protein structure space, and its implications for protein structure prediction are discussed. (C) 2003 Elsevier Ltd. All rights reserved.