Enzyme function less conserved than anticipated

Enzyme function less conserved than anticipated
复制标题

DOI:
10.1016/s0022-2836(02)00016-5
复制
发表时间:
2002-04-26
影响因子:
5.6
通讯作者:
Rost, B
Rost, B
中科院分区:
生物学2区
文献类型:
--
作者:
Rost, B

文献摘要

被引文献

相似文献

暗示蛋白质结构相似性的序列相似性水平已得到很好的确立。最近,许多研究小组提出了序列相似性的阈值,这意味着酶功能的相似性。所有先前的结果表明酶功能的强保守性高于50%成对序列同一性的水平。在这里,我认为,所有的团体大大高估了酶功能的保守性,因为他们的数据集要么太有偏见,或太小。无偏分析表明,低于30%的对片段的50%以上的序列同一性具有完全相同的EC数。另一个令人惊讶的发现是,即使低于10(-50)的BLAST E值也不足以无错误地自动转移酶功能。正如预期的那样,大多数错误分类起源于相对较短区域的相似性和/或转移不同域的注释。这两个问题都不能通过调整基因组注释自动转移的阈值来容易地纠正。对于高序列相似性,将序列同一性与比对长度(与HSSP阈值的距离)相关的评分优于统计BLAST评分。特别地,距离分数允许10%最相似的酶对的酶功能的无错误转移。这些结果说明了评估蛋白质功能的保守性和保证基因组注释无误是多么困难,一般来说:具有数百万对比较的集合可能不足以得出统计学上显著的结论。在实践中,修订后的详细估计酶功能的序列保守性可能会提供重要的基准,为日常序列分析和更谨慎的自动基因组注释。(C)2002爱思唯尔科技有限公司版权所有。
The level of sequence similarity that implies similarity in protein structure is well established. Recently, many groups proposed thresholds for similarity in sequence implying similarity in enzymatic function. All previous results suggest the strong conservation of enzymatic function above levels of 50% pairwise sequence identity. Here, I argue that all groups substantially overestimated the conservation of enzyme function because their data sets were either too biased, or too small. An unbiased analysis suggested that less than 30% of the pair fragments above 50% sequence identity have entirely identical EC numbers. Another surprising finding was that even BLAST E-values below 10(-50) did not suffice to automatically transfer enzyme function without errors. As expected, most misclassifications originated from similarities in relatively short regions and/or from transferring annotations for different domains. Both problems cannot be corrected easily by adjusting the thresholds for automatic transfer of genome annotations. A score relating sequence identity to alignment length (distance from HSSP-threshold) outperformed statistical BLAST scores for high sequence similarity. In particular, the distance score allowed error-free transfer of enzyme function for the 10% most similar enzyme pairs. The results illustrated how difficult it is to assess the conservation of protein function and to guarantee error-free genome annotations, in general: sets with millions of pair comparisons might not suffice to arrive at statistically significant conclusions. In practice, the revised detailed estimates for the sequence conservation of enzyme function may provide important benchmarks for everyday sequence analysis and for more cautious automatic genome annotations. (C) 2002 Elsevier Science Ltd. All rights reserved.