The intrinsic dimension of protein sequence evolution

The intrinsic dimension of protein sequence evolution
复制标题

DOI:
10.1371/journal.pcbi.1006767
复制
发表时间:
2019-04-01
影响因子:
4.3
通讯作者:
Laio, Alessandro
Laio, Alessandro
中科院分区:
生物学2区
文献类型:
--
作者:
Facco, Elena;Pagnani, Andrea;Laio, Alessandro

文献摘要

被引文献

相似文献

众所周知,为了保持其结构和功能,蛋白质不能随机改变其序列,而只能通过优先发生在特定位置的突变。在这里,我们定量研究蛋白质序列进化中允许的变异量,通过计算属于一个选择的蛋白质家族的序列的内在维度(ID)。ID是从给定序列开始的进化可以采取的独立方向的数量的度量。我们发现,ID几乎是恒定的序列属于同一个家庭,而且它是非常相似的,在不同的家庭,与6和12之间的值。这些值明显小于氨基酸的原始数量,证实了不同位点突变之间相关性的重要性。然而,我们证明,相关性是不足以解释我们在蛋白质家族中观察到的ID的小值。事实上,我们表明,ID的一组蛋白质序列产生的最大熵模型,一种方法,其中相关性占,通常显着大于在天然蛋白质家族中观察到的值。我们进一步证明了复制自然ID的一个关键因素是考虑序列的进化。作者概要蛋白质序列进化是一个极其复杂的过程,其作用最终取决于生物体适应环境变化的必要性。我们在这里解决一个基本的问题与这个过程有关:在多少个独立的方向可以演变的序列,而不损害蛋白质的折叠能力和执行其功能?我们发现,这些方向的数量是令人惊讶的小,在我们考虑的大多数家庭的10或更少。我们考虑的大多数理论模型都没有正确地解释这种性质,这些模型预测序列演化可以在30-40个独立的方向上发生。完成生成低维序列的任务的唯一方法是考虑序列同源性。
It is well known that, in order to preserve its structure and function, a protein cannot change its sequence at random, but only by mutations occurring preferentially at specific locations. We here investigate quantitatively the amount of variability that is allowed in protein sequence evolution, by computing the intrinsic dimension (ID) of the sequences belonging to a selection of protein families. The ID is a measure of the number of independent directions that evolution can take starting from a given sequence. We find that the ID is practically constant for sequences belonging to the same family, and moreover it is very similar in different families, with values ranging between 6 and 12. These values are significantly smaller than the raw number of amino acids, confirming the importance of correlations between mutations in different sites. However, we demonstrate that correlations are not sufficient to explain the small value of the ID we observe in protein families. Indeed, we show that the ID of a set of protein sequences generated by maximum entropy models, an approach in which correlations are accounted for, is typically significantly larger than the value observed in natural protein families. We further prove that a critical factor to reproduce the natural ID is to take into consideration the phylogeny of sequences.Author summary Protein sequence evolution is an extremely complex process, whose roles are ultimately determined by the necessity of living organisms to adapt to changes in the environment. We here address a fundamental question related with this process: in how many independent directions can a sequence evolve, without compromising the protein capability of folding and of performing its function? We find that the number of these directions is surprisingly small, of 10 or less in most of the families we considered. This property is not correctly accounted for by most of the theoretical model we considered, which predict that sequence evolution can take place in 30-40 independent directions. The only way to accomplish the task of generating low-dimensional sequences is to take into consideration sequence phylogeny.