Flavors of protein disorder

Flavors of protein disorder
复制标题

DOI:
10.1002/prot.10437
复制
发表时间:
2003-09-01
期刊:
PROTEINS-STRUCTURE FUNCTION AND GENETICS
影响因子:
--
通讯作者:
Obradovic, Z
Obradovic, Z
中科院分区:
其他
文献类型:
--
作者:
Vucetic, S;Brown, CJ;Obradovic, Z

文献摘要

被引文献

相似文献

内含子无序蛋白质的特征是在其天然状态下缺乏3-D结构的长区域,然而迄今为止它们与28种可区分的功能相关。以前的研究表明,从一种类型的蛋白质的无序训练的蛋白质预测往往实现差的准确性对不同类型的蛋白质的无序,从而表明无序蛋白质之间的序列特性的显着差异。重要的生物学问题是识别不同类型或风味的疾病,并检查它们与蛋白质功能的关系。由于相对缺乏与蛋白质紊乱相关的实验数据和背景知识,需要创新地使用计算方法来解决这些问题。我们开发了一种算法,该算法基于越来越多的预测因子之间的竞争将蛋白质紊乱划分为风味,预测准确性决定了不同预测因子的数量和单个蛋白质的划分。利用145个具有长(>30个氨基酸)无序区的不同特征的蛋白质,通过该方法鉴定了3种风味,称为V、C和S,其中V子集包含52个片段和7743个残基,C包含39个片段和3402个残基,S包含54个片段和5752个残基。V、C和S风味可通过氨基酸组成、序列位置和生物学功能来区分。对于SwissProt和28个基因组中的序列,它们的蛋白质功能表现出与不同紊乱风味的共性和使用的相关性,这表明这些蛋白质组中存在不同的风味功能集。总之,本文的结果支持风味功能的方法作为一个有用的补充结构基因组学作为一种手段,用于自动分配可能的功能序列。(C)2003 Wiley-Liss,Inc.
Intrinsically disordered proteins are characterized by long regions lacking 3-D structure in their native states, yet they have been so far associated with 28 distinguishable functions. Previous studies showed that protein predictors trained on disorder from one type of protein often achieve poor accuracy on disorder of proteins of a different type, thus indicating significant differences in sequence properties among disordered proteins. Important biological problems are identifying different types, or flavors, of disorder and examining their relationships with protein function. Innovative use of computational methods is needed in addressing these problems due to relative scarcity of experimental data and background knowledge related to protein disorder. We developed an algorithm that partitions protein disorder into flavors based on competition among increasing numbers of predictors, with prediction accuracy determining both the number of distinct predictors and the partitioning of the individual proteins. Using 145 variously characterized proteins with long (>30 amino acids) disordered regions, 3 flavors, called V, C, and S, were identified by this approach, with the V subset containing 52 segments and 7743 residues, C containing 39 segments and 3402 residues, and S containing 54 segments and 5752 residues. The V, C, and S flavors were distinguishable by amino acid compositions, sequence locations, and biological function. For the sequences in SwissProt and 28 genomes, their protein functions exhibit correlations with the commonness and usage of different disorder flavors, suggesting different flavor-function sets across these protein groups. Overall, the results herein support the flavor-function approach as a useful complement to structural genomics as a means for automatically assigning possible functions to sequences. (C) 2003 Wiley-Liss, Inc.