Can sequence-specific and dynamics-based metrics allow us to decipher the function in IDP sequences?

Can sequence-specific and dynamics-based metrics allow us to decipher the function in IDP sequences?
复制标题

DOI:
10.1016/j.bpj.2021.04.008
复制
发表时间:
2021-04
影响因子:
3.4
通讯作者:
S. Ozkan
S. Ozkan
中科院分区:
生物学3区
文献类型:
--
作者:
S. Ozkan

文献摘要

相似文献

在地球上所有生物体共有的生物分子中,蛋白质具有一系列基于核苷酸的对应物所不具有的功能。蛋白质之所以成为大自然的奇迹,是因为它们特定的氨基酸序列,编码了它们所有的生物物理特性。传统上,序列功能关系是通过蛋白质天然状态的结构表征来解码的。然而,并非所有蛋白质序列都需要其天然(功能)状态的结构。这种结构独立的蛋白质,也称为本质无序蛋白质 (IDP),在许多生物功能中至关重要,并且构成蛋白质组的相当一部分 (1-3)。尽管进行了二十年的广泛研究,但由于经典序列-结构-功能映射的失败,破译序列-功能代码仍然具有挑战性 (3)。在未结合状态下缺乏有序的紧凑形式,并且在结合形式中表现出结构混乱,这给结构对齐带来了挑战 (2)。此外,IDP 似乎比结构蛋白进化得更快,因为氨基酸插入和删除在 IDP 中更常见,从而使序列比对变得复杂 (3)。将 IDP 和结构蛋白统一起来的是功能状态(即结构蛋白的天然状态)下的构象整体以及相关的平衡动力学(图 1 A)。这个由一维序列决定的系综是该函数的基础。对于结构蛋白,通过全原子模拟或实验确定的三维相互作用的粗粒度模型对天然状态整体进行采样,使我们能够破译序列函数范式 (4)。然而,由于构象空间很大,这个过程对于 IDP 来说更具挑战性。尽管成功开发了粗粒模型并修改了全原子力场以采样 IDP 构象 (5),但构象系综的准确表示仍然存在争议 (6, 7)。因此,迫切需要基于序列的第一原理理论模型来描述 IDP 的整体特征,以破译其序列编码功能。
Among the biomolecules common to all living organisms on Earth, proteins conduct a diverse array of functions that their nucleotide-based counterparts do not. What makes proteins nature’s miracle is their specific amino acid sequence, which encodes all their biophysical properties. Traditionally, sequence-function relation is decoded by structural characterization of proteins in their native states. However, not all protein sequences require a structure in their native (functional) state. Such structure-independent proteins, also known as intrinsically disordered proteins (IDPs), are critical in many biological functions and constitute a considerable fraction of the proteome (1–3). Despite two decades of extensive studies, it remains challenging to decipher sequence-function code because of failures in classical sequence-structure-function mapping (3). Lacking an ordered compact form in the unbound state along with exhibiting structural promiscuity in bound forms brings challenges in structural alignments (2). Furthermore, IDPs appear to evolve faster than structured proteins because insertion and deletion of amino acids are more common in IDPs, thereby complicating sequence alignments (3).What unifies IDPs and structured proteins is the conformational ensemble in the functional state (ie, native state for structural proteins) and associated equilibrium dynamics (Fig. 1 A). This ensemble, dictated by the one-dimensional sequence, underlies the function. For structural proteins, sampling a native-state ensemble through all-atom simulations or coarse-grain models of experimentally determined threedimensional interactions allows us to decipher the sequence-function paradigm (4). However, this process is much more challenging for IDPs because of large conformational space. Despite the success of development of coarse-grain models and modifications in all-atom forcefields to sample IDP conformations (5), accurate representation of the conformational ensemble is still under debate (6, 7). Thus, sequence-based first-principle theoretical models that describe ensemble features of IDPs are desperately needed to decipher their sequence-encoded functions.