Sequence-similar, structure-dissimilar protein pairs in the PDB.

Sequence-similar, structure-dissimilar protein pairs in the PDB.
复制标题

DOI:
10.1002/prot.21770
复制
发表时间:
2008-05-01
影响因子:
2.9
通讯作者:
Kolodny, Rachel
Kolodny, Rachel
中科院分区:
生物学4区
文献类型:
--
作者:
Kosloff, Mickey;Kolodny, Rachel

文献摘要

参考文献

被引文献

相似文献

人们通常认为,在蛋白质数据库(PDB)中,具有相似序列的两种蛋白质也将具有相似的结构。因此,事实证明,基于基于序列的相似性标准,开发已删除“冗余”结构的 PDB 子集是有用的。类似地,当使用同源建模预测蛋白质结构时,如果仅通过序列选择用于建模目标序列的模板结构,则这隐含地假设所有序列相似的模板是等效的。在这里,我们表明这种假设通常是不正确的,创建 PDB 子集的标准方法可能会导致结构和功能上重要信息的丢失。我们对大量蛋白质对进行了基于序列的结构叠加和基于几何的结构比对,以确定序列相似性在多大程度上保证了结构相似性。我们发现许多例子,其中两种序列相似的蛋白质具有彼此显着不同的结构。结构差异的根源通常具有功能基础。所识别的此类蛋白质对的数量以及差异的程度取决于用于计算差异的方法;特别是基于序列的结构叠加将比基于几何的结构比对识别更多数量的结构不同对。当两个序列可以以统计上有意义的方式比对时,基于序列的结构叠加提供了结构差异的有意义的测量。这种方法和基于几何的结构对齐揭示了一些不同的信息,并且在给定的应用中,其中一种或另一种可能更可取。我们的结果表明,在某些情况下,特别是同源建模,根据序列从 PDB 中挑选的非冗余数据集的常见使用可能会掩盖重要的结构和功能信息。我们建立了一个序列相似、结构不同的蛋白质对的数据库,这将有助于解决这个问题(http://luna.bioc.columbia.edu/rachel/seqsimstrdiff.htm)。
It is often assumed that in the Protein Data Bank (PDB), two proteins with similar sequences will also have similar structures. Accordingly, it has proved useful to develop subsets of the PDB from which “redundant” structures have been removed, based on a sequence-based criterion for similarity. Similarly, when predicting protein structure using homology modeling, if a template structure for modeling a target sequence is selected by sequence alone, this implicitly assumes that all sequence-similar templates are equivalent. Here, we show that this assumption is often not correct and that standard approaches to create subsets of the PDB can lead to the loss of structurally and functionally important information. We have carried out sequence-based structural superpositions and geometry-based structural alignments of a large number of protein pairs to determine the extent to which sequence similarity ensures structural similarity. We find many examples where two proteins that are similar in sequence have structures that differ significantly from one another. The source of the structural differences usually has a functional basis. The number of such proteins pairs that are identified and the magnitude of the dissimilarity depend on the approach that is used to calculate the differences; in particular sequence-based structure superpositioning will identify a larger number of structurally dissimilar pairs than geometry-based structural alignments. When two sequences can be aligned in a statistically meaningful way, sequence-based structural superpositioning provides a meaningful measure of structural differences. This approach and geometry-based structure alignments reveal somewhat different information and one or the other might be preferable in a given application. Our results suggest that in some cases, notably homology modeling, the common use of nonredundant datasets, culled from the PDB based on sequence, may mask important structural and functional information. We have established a data base of sequence-similar, structurally dissimilar protein pairs that will help address this problem (http://luna.bioc.columbia.edu/rachel/seqsimstrdiff.htm).
DOI: 10.1016/j.jmb.2005.05.066
发表时间: 2005-08-12
影响因子: 5.6
作者:
Eyal, E;Gerzon, S;Sobolev, V
通讯作者: Sobolev, V
DOI: 10.1371/journal.pcbi.0020155
发表时间: 2006-11-17
影响因子: 4.3
作者:
Levy ED;Pereira-Leal JB;Chothia C;Teichmann SA
通讯作者: Teichmann SA
DOI: 10.1002/prot.10168
发表时间: 2002-09-01
影响因子: 2.9
作者:
Krebs, WG;Alexandrov, V;Gerstein, M
通讯作者: Gerstein, M
DOI: 10.1074/jbc.m411155200
发表时间: 2005-01-28
影响因子: 4.8
作者:
Ködding, J;Killig, F;Welte, W
通讯作者: Welte, W
DOI: 10.1093/nar/gkh039
发表时间: 2004-01-01
影响因子: 14.9
作者:
Andreeva, A;Howorth, D;Murzin, AG
通讯作者: Murzin, AG