Structural footprinting in protein structure comparison: the impact of structural fragments.

Structural footprinting in protein structure comparison: the impact of structural fragments.
复制标题

DOI:
10.1186/1472-6807-7-53
复制
发表时间:
2007-08-09
影响因子:
--
通讯作者:
Przytycka TM
Przytycka TM
中科院分区:
生物4区
文献类型:
--
作者:
Zotenko E;Dogan RI;Wilbur WJ;O'Leary DP;Przytycka TM

文献摘要

参考文献

相似文献

加速蛋白质结构比较的一种方法是投影方法,其中将蛋白质结构映射到高维向量,并通过相应向量之间的距离来近似结构相似性。结构足迹方法是采用相同通用技术来生成映射的投影方法:首先选择一组代表性的结构片段作为模型,然后将蛋白质结构映射到向量,其中每个维度对应于特定模型,并“计算”模型在结构中出现的次数。任何两种结构足迹方法之间的主要区别在于它们使用的模型集;事实上,通过改变所使用的结构片段的类型及其表示的细节量,可以生成大量方法。这些选择如何影响该方法检测各种类型结构相似性的能力?为了回答这个问题,我们对三种结构足迹方法进行了基准测试,这些方法在模型选择方面与 CATH 数据库存在很大差异。在第一组实验中,我们比较了这些方法检测进化相关结构(即同一 CATH 超家族内的结构)的结构相似性特征的能力。在第二组实验中,我们测试了这些方法与 CATH 层次结构的类、体系结构和折叠级别的分类组所施加的边界的一致性。在这两个实验中,我们发现使用二级结构信息的方法平均具有最佳性能,但没有一种方法在给定分类级别的所有组中始终表现最佳。我们还发现,结合这些方法的输出可以显着提高性能。此外,我们用于测量和可视化方法与 CATH 层次结构的一致性的新技术(包括阈值亲和力图)在这项工作之外也很有用。特别是,它们可用于根据该方法使用的结构片段来揭示不同分类组的相似组成,从而提供蛋白质结构宇宙的连续性质的替代证明。
One approach for speeding-up protein structure comparison is the projection approach, where a protein structure is mapped to a high-dimensional vector and structural similarity is approximated by distance between the corresponding vectors. Structural footprinting methods are projection methods that employ the same general technique to produce the mapping: first select a representative set of structural fragments as models and then map a protein structure to a vector in which each dimension corresponds to a particular model and "counts" the number of times the model appears in the structure. The main difference between any two structural footprinting methods is in the set of models they use; in fact a large number of methods can be generated by varying the type of structural fragments used and the amount of detail in their representation. How do these choices affect the ability of the method to detect various types of structural similarity? To answer this question we benchmarked three structural footprinting methods that vary significantly in their selection of models against the CATH database. In the first set of experiments we compared the methods' ability to detect structural similarity characteristic of evolutionarily related structures, i.e., structures within the same CATH superfamily. In the second set of experiments we tested the methods' agreement with the boundaries imposed by classification groups at the Class, Architecture, and Fold levels of the CATH hierarchy. In both experiments we found that the method which uses secondary structure information has the best performance on average, but no one method performs consistently the best across all groups at a given classification level. We also found that combining the methods' outputs significantly improves the performance. Moreover, our new techniques to measure and visualize the methods' agreement with the CATH hierarchy, including the threshholded affinity graph, are useful beyond this work. In particular, they can be used to expose a similar composition of different classification groups in terms of structural fragments used by the method and thus provide an alternative demonstration of the continuous nature of the protein structure universe.
DOI: 10.1016/s0097-8485(96)80004-0
发表时间: 1996-03-01
期刊: COMPUTERS & CHEMISTRY
影响因子: --
作者:
Gribskov, M;Robinson, NL
通讯作者: Robinson, NL
DOI: 10.1073/pnas.2636460100
发表时间: 2003-01-07
影响因子: 11.1
作者:
Rogen, P;Fain, B
通讯作者: Fain, B
DOI: 10.1073/pnas.68.4.815
发表时间: 1971-01-01
影响因子: 11.1
作者:
FULLER, FB
通讯作者: FULLER, FB
DOI: 10.1006/jmbi.2001.5250
发表时间: 2002-01-25
影响因子: 5.6
作者:
Carugo, O;Pongor, S
通讯作者: Pongor, S
DOI: 10.1073/pnas.0308656100
发表时间: 2004-03-16
影响因子: 11.1
作者:
Choi, IG;Kwon, J;Kim, SH
通讯作者: Kim, SH