Target space for structural genomics revisited

Target space for structural genomics revisited
复制标题

DOI:
10.1093/bioinformatics/18.7.922
复制
发表时间:
2002-07-01
期刊:
影响因子:
5.8
通讯作者:
Rost, B
Rost, B
中科院分区:
生物学3区
文献类型:
--
作者:
Liu, JF;Rost, B

文献摘要

被引文献

相似文献

动机:结构基因组学最终旨在确定所有蛋白质的结构。然而,在开始时,实验者可能会专注于球状蛋白质,以实现蛋白质序列空间的快速基本覆盖。结构基因组学要针对多少蛋白质?有多少蛋白质将被排除在外,因为我们已经有了这些蛋白质的结构信息,或者因为它们不是球形的?我们必须回答这些问题的背景下,我们的目标选择的东北结构基因组学联盟(NESG)。结果:我们估计,结构信息可用于约6-38%的所有蛋白质; 6%,如果我们需要高精度的比较建模,38%,如果我们满意有一个粗略的想法折叠。排除所有非球形区域,我们发现结构基因组学可能必须针对所有蛋白质的约48%。这对应于整个蛋白质组的残基百分比相似(52%)。我们探索了许多不同的策略来聚类蛋白质空间,以找到代表这48%结构未知蛋白质的家族的数量。对于所有完全测序的真核生物的子集,我们发现了超过18000个片段簇,每个片段簇都可能是结构基因组学的合适靶点。
Motivation: Structural genomics eventually aims at determining structures for all proteins. However, in the beginning experimentalists are likely to focus on globular proteins to achieve a rapid basic coverage of protein sequence space. How many proteins will structural genomics have to target? How many proteins will be excluded since we already have structural information for these or since they are not globular? We have to answer these questions in the context of our target selection for the North-East Structural Genomics Consortium (NESG).Results: We estimated that structural information is available for about 6-38% of all proteins; 6% if we require high accuracy in comparative modelling, 38% if we are satisfied with having a rough idea about the fold. Excluding all regions that are not globular, we found that structural genomics may have to target about 48% of all proteins. This corresponded to a similar percentage of residues of the entire proteomes (52%). We explored a number of different strategies to cluster protein space in order to find the number of families representing these 48% of structurally unknown proteins. For the subset of all entirely sequenced eukaryotes, we found over 18 000 fragment clusters each of which may be a suitable target for structural genomics.