Prediction of functional sites in proteins using conserved functional group analysis

Prediction of functional sites in proteins using conserved functional group analysis
复制标题

DOI:
10.1016/j.jmb.2004.01.053
复制
发表时间:
2004-04-02
影响因子:
5.6
通讯作者:
Sowdhamini, R
Sowdhamini, R
中科院分区:
生物学2区
文献类型:
--
作者:
Innis, CA;Anand, AP;Sowdhamini, R

文献摘要

被引文献

相似文献

对蛋白质功能位点的详细了解是在分子水平上了解其作用模式的绝对先决条件。然而,蛋白质序列和结构信息积累的快速速度远远超出了我们通过实验确定其生化作用的能力。因此,需要能够有效处理大量数据中包含的进化信息的计算方法,特别是与功能上重要的位点和残基的性质和位置相关的信息。这里介绍的方法称为保守功能基团 (CFG) 分析,依赖于氨基酸侧链中发现的化学基团的简化表示,以从单个蛋白质结构及其许多序列同源物中识别功能位点。我们表明,CFG 分析可以完全或部分预测功能位点的位置,与 470 个测试病例中的 96% 类似,并且与其他可用方法不同,它能够容忍序列同一性的广泛变化。此外,我们还讨论了它在结构基因组学背景下的潜力,其中自动化、可扩展性和效率至关重要,并且越来越多的蛋白质结构是在没有功能先验知识的情况下确定的。我们对假设蛋白质 Ydde_Ecoli 的分析证明了这一点,该蛋白质的结构最近由东北结构基因组学联盟的成员解决。尽管该蛋白质的拟议活性位点需要通过实验进行验证,但此示例说明了 CFG 分析作为识别可能在蛋白质生化功能中发挥重要作用的残基的通用工具的范围。因此,我们的方法为结构基因组学项目提供了一种方便的解决方案。 (C) 2004 Elsevier Ltd. 保留所有权利。
A detailed knowledge of a protein's functional site is an absolute prerequisite for understanding its mode of action at the molecular level. However, the rapid pace at which sequence and structural information is being accumulated for proteins greatly exceeds our ability to determine their biochemical roles experimentally. As a result, computational methods are required which allow for the efficient processing of the evolutionary information contained in this wealth of data, in particular that related to the nature and location of functionally important sites and residues. The method presented here, referred to as conserved functional group (CFG) analysis, relies on a simplified representation of the chemical groups found in amino acid side-chains to identify functional sites from a single protein structure and a number of its sequence homologues. We show that CFG analysis can fully or partially predict the location of functional sites in similar to96% of the 470 cases tested and that, unlike other methods available, it is able to tolerate wide variations in sequence identity. In addition, we discuss its potential in a structural genomics context, where automation, scalability and efficiency are critical, and an increasing number of protein structures are determined with no prior knowledge of function. This is exemplified by our analysis of the hypothetical protein Ydde_Ecoli, whose structure was recently solved by members of the North East Structural Genomics consortium. Although the proposed active site for this protein needs to be validated experimentally, this example illustrates the scope of CFG analysis as a general tool for the identification of residues likely to play an important role in a protein's biochemical function. Thus, our method offers a convenient solution to to emerge from structural genomics projects. (C) 2004 Elsevier Ltd. All rights reserved.