Amino acid positions subject to multiple coevolutionary constraints can be robustly identified by their eigenvector network centrality scores.

Amino acid positions subject to multiple coevolutionary constraints can be robustly identified by their eigenvector network centrality scores.
复制标题

DOI:
10.1002/prot.24948
复制
发表时间:
2015-12
期刊:
影响因子:
2.9
通讯作者:
Swint-Kruse L
Swint-Kruse L
中科院分区:
生物学4区
文献类型:
--
作者:
Parente DJ;Ray JC;Swint-Kruse L

文献摘要

被引文献

相似文献

随着蛋白质的进化,对蛋白质结构或功能至关重要的氨基酸位置受到突变的限制。这些位置可以通过分析序列家族的氨基酸保守性或位置对之间的共进化来检测。协同进化的得分通常是排序和阈值,以揭示最高的成对得分,但它们也可以被视为加权网络。在这里,我们使用网络分析来绕过共同进化研究的一个主要复杂性:对于给定的序列比对,替代算法通常会识别不同的、最高的成对分数。我们调和的结果,从五个常用的,数学上不同的算法(ELSC,McBASC,OMES,SCA和ZNMI),使用的LacI/GalR和1,6-二磷酸醛缩酶蛋白质家族作为模型。计算使用未阈值的共同进化分数,从中减去列特定的属性,如序列熵和随机噪声;通过计算各种网络中心性分数来确定“中心”位置。当比较算法时,网络中心性方法,特别是特征向量中心性方法,比最高成对分数的比较显示出明显更好的一致性。具有大中心性分数的位置发生在关键结构位置和/或对突变功能敏感。此外,最高的中心位置通常与那些具有最高成对共同进化分数的位置不同:中心位置通常具有多个中等分数,而不是几个强分数。我们的结论是,特征向量中心计算揭示了一个强大的进化模式的约束-检测不同的算法-发生在关键蛋白质的位置。最后,我们讨论的事实是,多种模式共存的进化数据,一起,引起紧急蛋白质功能。
As proteins evolve, amino acid positions key to protein structure or function are subject to mutational constraints. These positions can be detected by analyzing sequence families for amino acid conservation or for co-evolution between pairs of positions. Co-evolutionary scores are usually rank-ordered and thresholded to reveal the top pairwise scores, but they also can be treated as weighted networks. Here, we used network analyses to bypass a major complication of co-evolution studies: For a given sequence alignment, alternative algorithms usually identify different, top pairwise scores. We reconciled results from five commonly-used, mathematically divergent algorithms (ELSC, McBASC, OMES, SCA, and ZNMI), using the LacI/GalR and 1,6-bisphosphate aldolase protein families as models. Calculations used unthresholded co-evolution scores from which column-specific properties such as sequence entropy and random noise were subtracted; “central” positions were identified by calculating various network centrality scores. When compared among algorithms, network centrality methods, particularly eigenvector centrality, showed markedly better agreement than comparisons of the top pairwise scores. Positions with large centrality scores occurred at key structural locations and/or were functionally sensitive to mutations. Further, the top central positions often differed from those with top pairwise co-evolution scores: Instead of a few strong scores, central positions often had multiple, moderate scores. We conclude that eigenvector centrality calculations reveal a robust evolutionary pattern of constraints – detectable by divergent algorithms – that occur at key protein locations. Finally, we discuss the fact that multiple patterns co-exist in evolutionary data that, together, give rise to emergent protein functions.