Protein intrinsically disordered regions have a non-random, modular architecture.

Protein intrinsically disordered regions have a non-random, modular architecture.
复制标题

DOI:
10.1093/bioinformatics/btad732
复制
发表时间:
2023-12-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
通讯作者:
--
中科院分区:
其他
文献类型:
--
作者:

文献摘要

参考文献

相似文献

蛋白质序列可以大致分为两类:采用稳定二级结构并折叠成结构域的蛋白质(即球状蛋白质),以及不采用稳定二级结构并折叠成结构域的蛋白质。属于后一类的序列在构象上是不均匀的,并且被描述为本质上是无序的。几十年来对球状蛋白质结构和功能的研究已经产生了一套计算工具,可以根据结构域类型对其进行细分,这种方法彻底改变了我们理解和预测蛋白质功能的方式。相反地,不知道无序蛋白质区域的序列是否服从于将使它们能够进行亚分类的广泛可概括的组织原则。在这里,我们报告了一种统计方法的发展,该方法可以量化整个序列中氨基酸组成的线性方差。多个例子,我们提供的证据表明,本质上无序的区域组织成统计非随机模块的独特的成分偏见。模块化观察到低和高复杂度的序列,在某些情况下,我们发现,模块组织在重复的模式。这些数据表明,无序序列是非随机组织成模块化架构,并激励未来的实验,全面分类模块类型,并确定模块构成功能上可分离的单位类似的球状蛋白质的结构域的程度。可在https://github.com/MWPlabUTSW/Chi-Score-Analysis.git上免费获得复制所有图的源代码、文档和数据。该分析也可作为Google Colab Notebook(https://colab.research.google.com/github/MWPlabUTSW/Chi-Score-Analysis/blob/main/ChiScore_Analysis.ipynb)获得。
Protein sequences can be broadly categorized into two classes: those which adopt stable secondary structure and fold into a domain (i.e. globular proteins), and those that do not. The sequences belonging to this latter class are conformationally heterogeneous and are described as being intrinsically disordered. Decades of investigation into the structure and function of globular proteins has resulted in a suite of computational tools that enable their sub-classification by domain type, an approach that has revolutionized how we understand and predict protein functionality. Conversely, it is unknown if sequences of disordered protein regions are subject to broadly generalizable organizational principles that would enable their sub-classification. Here, we report the development of a statistical approach that quantifies linear variance in amino acid composition across a sequence. With multiple examples, we provide evidence that intrinsically disordered regions are organized into statistically non-random modules of unique compositional bias. Modularity is observed for both low and high-complexity sequences and, in some cases, we find that modules are organized in repetitive patterns. These data demonstrate that disordered sequences are non-randomly organized into modular architectures and motivate future experiments to comprehensively classify module types and to determine the degree to which modules constitute functionally separable units analogous to the domains of globular proteins. The source code, documentation, and data to reproduce all figures are freely available at https://github.com/MWPlabUTSW/Chi-Score-Analysis.git. The analysis is also available as a Google Colab Notebook (https://colab.research.google.com/github/MWPlabUTSW/Chi-Score-Analysis/blob/main/ChiScore_Analysis.ipynb).
DOI: 10.7554/elife.77058
发表时间: 2022-09-13
期刊: ELIFE
影响因子: 7.7
作者:
Lee, Byron;Jaberi-Lashkari, Nima;Calo, Eliezer
通讯作者: Calo, Eliezer
DOI: 10.1038/s41586-021-03819-2
发表时间: 2021-08
期刊: Nature
影响因子: 64.8
作者:
Jumper J;Evans R;Pritzel A;Green T;Figurnov M;Ronneberger O;Tunyasuvunakool K;Bates R;Žídek A;Potapenko A;Bridgland A;Meyer C;Kohl SAA;Ballard AJ;Cowie A;Romera-Paredes B;Nikolov S;Jain R;Adler J;Back T;Petersen S;Reiman D;Clancy E;Zielinski M;Steinegger M;Pacholska M;Berghammer T;Bodenstein S;Silver D;Vinyals O;Senior AW;Kavukcuoglu K;Kohli P;Hassabis D
通讯作者: Hassabis D
DOI: 10.1007/bf02703118
发表时间: 1993-06-01
影响因子: 2.9
作者:
MITRA, CK;RANI, M
通讯作者: RANI, M
DOI: 10.1038/s41594-019-0248-4
发表时间: 2019-07-01
影响因子: 16.8
作者:
Cao, Qin;Boyer, David R.;Eisenberg, David S.
通讯作者: Eisenberg, David S.
DOI: 10.1093/bioinformatics/btaa1045
发表时间: 2020-12-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Necci, Marco;Piovesan, Damiano;Tosatto, Silvio C. E.
通讯作者: Tosatto, Silvio C. E.