Analysis and prediction of functional sub-types from protein sequence alignments

Analysis and prediction of functional sub-types from protein sequence alignments
复制标题

DOI:
10.1006/jmbi.2000.4036
复制
发表时间:
2000-10-13
影响因子:
5.6
通讯作者:
Russell, RB
Russell, RB
中科院分区:
生物学2区
文献类型:
--
作者:
Hannenhalli, SS;Russell, RB

文献摘要

被引文献

相似文献

蛋白质序列家族的数量和多样性不断增加,需要新的方法来定义和预测有关功能的细节。在这里,我们提出了一种从多个蛋白质序列比对中分析和预测功能亚型的方法。给定比对和根据某些功能定义(例如酶特异性)分组为子类型的蛋白质组,该方法通过比较子类型特异性序列概况和分析比对中的位置熵来识别指示功能差异的位置。具有显着高位置相对熵的比对位置与已知参与定义核苷酸环化酶、蛋白激酶、乳酸/苹果酸脱氢酶和胰蛋白酶样丝氨酸蛋白酶的亚型的那些相关。我们强调了这些蛋白质的新位​​置,建议进行额外的实验来阐明特异性的基础。该方法还能够预测未分类序列的子类型。我们评估了预测方法的几种变体,并将它们与简单的序列比较进行比较。为了进行评估,我们删除了与要进行预测的序列的密切同源物(通过高于阈值的序列同一性)。这模拟了已知蛋白质属于蛋白质家族但不是已知亚型的另一种蛋白质的近亲的情况。考虑到上述四个家族,以及 30% 的序列同一性阈值,我们的最佳方法给出的准确度为 96%,而序列相似性为 80%,BLAST 为 74%。我们描述了从 PFAM 和 SWISSPROT 数据库中自动解析比对得出的一组子类型分组的推导,并使用它来执行大规模评估。最佳方法的平均准确度为 94%,而序列相似性的平均准确度为 68%,BLAST 的平均准确度为 79%。我们讨论了对实验设计、基因组注释以及蛋白质功能和蛋白质残基内距离的预测的影响。 (C) 2000 年学术出版社。
The increasing number and diversity of protein sequence families requires new methods to define and predict details regarding function. Here, we present a method for analysis and prediction of functional subtypes from multiple protein sequence alignments. Given an alignment and set of proteins grouped into sub-types according to some definition of function, such as enzymatic specificity, the method identifies positions that are indicative of functional differences by comparison of sub-type specific sequence profiles, and analysis of positional entropy in the alignment. Alignment positions with significantly high positional relative entropy correlate with those known to be involved in defining sub-types for nucleotidyl cyclases, protein kinases, lactate/malate dehydrogenases and trypsin-like serine proteases. We highlight new positions for these proteins that suggest additional experiments to elucidate the basis of specificity. The method is also able to predict sub-type for unclassified sequences. We assess several variations on a prediction method, and compare them to simple sequence comparisons. For assessment, we remove close homologues to the sequence for which a prediction is to be made (by a sequence identity above a threshold). This simulates situations where a protein is known to belong to a protein family, but is not a close relative of another protein of known sub-type. Considering the four families above, and a sequence identity threshold of 30 %, our best method gives an accuracy of 96% compared to 80% obtained for sequence similarity and 74% for BLAST. We describe the derivation of a set of sub-type groupings derived from an automated parsing of alignments from PFAM and the SWISSPROT database, and use this to perform a large-scale assessment. The best method gives an average accuracy of 94% compared to 68% for sequence similarity and 79% for BLAST. We discuss implications for experimental design, genome annotation and the prediction of protein function and protein intra-residue distances. (C) 2000 Academic Press.