Binary classification of protein molecules into intrinsically disordered and ordered segments.

Binary classification of protein molecules into intrinsically disordered and ordered segments.
复制标题

DOI:
10.1186/1472-6807-11-29
复制
发表时间:
2011-06-22
影响因子:
--
通讯作者:
Nishikawa K
Nishikawa K
中科院分区:
生物4区
文献类型:
--
作者:
Fukuchi S;Hosoda K;Homma K;Gojobori T;Nishikawa K

文献摘要

参考文献

被引文献

相似文献

尽管蛋白质的结构域(SDs)很重要,但目前人类蛋白质组中有一半的区域没有SD分配。这些未分配的区域不仅由新的SDs组成,而且也由内在无序(ID)区域组成,因为蛋白质,特别是真核生物中的蛋白质,通常含有相当一部分的ID区域。由于可以从氨基酸序列中推断出ID区域,因此结合SD和ID区域分配的方法可以确定任何蛋白质组中SD和ID区域的比例。与其他可用的ID预测程序仅仅识别可能的ID区域相比,我们之前开发的DICHOT系统将整个蛋白质序列分类为SDs和ID区域。DICHOT在人类蛋白质组中的应用表明,残基方向的ID区占35%,与PDB结构相似的SDs占52%,而与PDB结构不相似的SDs占13%。最后一组由新的结构域组成,称为隐结构域,是结构基因组学的良好靶点。DICHOT方法应用于其他模式生物的蛋白质组表明真核生物通常具有较高的ID含量,而原核生物则没有。在人类蛋白质中,亚细胞定位的ID含量不同:核蛋白的残基ID分数最高(47%),而线粒体蛋白的残基ID分数最低(13%)。磷酸化和o链糖基化位点被发现优先位于ID区域。由于o -链聚糖附着在蛋白质细胞外区域的残基上,这种修饰可能保护ID区域在细胞外环境中不受蛋白水解裂解的影响。选择性剪接事件往往更频繁地发生在ID区域。我们将此解释为自然选择在蛋白质水平上进行选择性剪接的证据。我们将整个蛋白质区域分为SDs和ID两类,从而获得各种全基因组统计数据。本研究结果是了解蛋白质结构结构的重要基础信息,并已在http://spock.genes.nig.ac.jp/~genome/DICHOT上公开发布。
Although structural domains in proteins (SDs) are important, half of the regions in the human proteome are currently left with no SD assignments. These unassigned regions consist not only of novel SDs, but also of intrinsically disordered (ID) regions since proteins, especially those in eukaryotes, generally contain a significant fraction of ID regions. As ID regions can be inferred from amino acid sequences, a method that combines SD and ID region assignments can determine the fractions of SDs and ID regions in any proteome. In contrast to other available ID prediction programs that merely identify likely ID regions, the DICHOT system we previously developed classifies the entire protein sequence into SDs and ID regions. Application of DICHOT to the human proteome revealed that residue-wise ID regions constitute 35%, SDs with similarity to PDB structures comprise 52%, while SDs with no similarity to PDB structures account for the remaining 13%. The last group consists of novel structural domains, termed cryptic domains, which serve as good targets of structural genomics. The DICHOT method applied to the proteomes of other model organisms indicated that eukaryotes generally have high ID contents, while prokaryotes do not. In human proteins, ID contents differ among subcellular localizations: nuclear proteins had the highest residue-wise ID fraction (47%), while mitochondrial proteins exhibited the lowest (13%). Phosphorylation and O-linked glycosylation sites were found to be located preferentially in ID regions. As O-linked glycans are attached to residues in the extracellular regions of proteins, the modification is likely to protect the ID regions from proteolytic cleavage in the extracellular environment. Alternative splicing events tend to occur more frequently in ID regions. We interpret this as evidence that natural selection is operating at the protein level in alternative splicing. We classified entire regions of proteins into the two categories, SDs and ID regions and thereby obtained various kinds of complete genome-wide statistics. The results of the present study are important basic information for understanding protein structural architectures and have been made publicly available at http://spock.genes.nig.ac.jp/~genome/DICHOT.
DOI: 10.1021/bi035934p
发表时间: 2004-03-23
期刊: BIOCHEMISTRY
影响因子: 2.9
作者:
Kumar, R;Betney, R;McEwan, IJ
通讯作者: McEwan, IJ
DOI: 10.1021/pr060049p
发表时间: 2006-04-01
影响因子: 4.4
作者:
Chen, JW;Romero, P;Dunker, AK
通讯作者: Dunker, AK
DOI: 10.1098/rstb.2002.1193
发表时间: 2003-01-29
影响因子: 6.3
作者:
Andersson, SGE;Karlberg, O;Kurland, CG
通讯作者: Kurland, CG
DOI: 10.1093/nar/gkp985
发表时间: 2010-01
影响因子: 14.9
作者:
Finn RD;Mistry J;Tate J;Coggill P;Heger A;Pollington JE;Gavin OL;Gunasekaran P;Ceric G;Forslund K;Holm L;Sonnhammer EL;Eddy SR;Bateman A
通讯作者: Bateman A
DOI: 10.1073/pnas.85.12.4335
发表时间: 1988-06-01
影响因子: 11.1
作者:
KOZARSKY, K;KINGSLEY, D;KRIEGER, M
通讯作者: KRIEGER, M