Functional coverage of the human genome by existing structures, structural genomics targets, and homology models.

Functional coverage of the human genome by existing structures, structural genomics targets, and homology models.
复制标题

现有结构、结构基因组目标和同源模型对人类基因组的功能覆盖。

DOI:
10.1371/journal.pcbi.0010031
复制
发表时间:
2005-08
影响因子:
4.3
通讯作者:
Bourne PE
Bourne PE
中科院分区:
生物学2区
文献类型:
--
作者:
Xie L;Bourne PE

文献摘要

参考文献

被引文献

相似文献

由于实验限制和结构生物学家对特定功能类蛋白质的定位而导致的蛋白质结构和功能空间的偏差早已被认识到,但从未连续量化。以酶委员会和基因本体论分类为参照框架,整合来自蛋白质数据库(PDB)的结构数据、来自结构基因组学项目的目标序列、来自超家族数据库的结构同源性以及来自EnSemb1和NCBI的基因组注释,我们在结构域和整个蛋白质水平上提供了相对于人类基因组的蛋白质结构和功能空间的当前和预测覆盖的量化视图。蛋白质结构目前提供了至少一个结构域,覆盖了基因组中所确定的37%的功能类别;25%的基因组存在整个结构覆盖。如果解决了所有的结构基因组学目标(是PDB中当前结构数量的两倍),估计一个结构域的结构将覆盖已确定的功能类的69%,完整结构覆盖率将为44%。现有实验结构的同源模型将37%的覆盖率扩展到56%的基因组作为单个结构域,25%到31%的完整结构。同源模型的覆盖率在蛋白质家族中并不均匀分布,反映了家族内部不同程度的序列和结构差异。虽然这些数据提供了覆盖范围,但反过来,它们也系统地突出了应该确定其结构的蛋白质的功能类别。这里突出显示了当前没有结构表示的关键功能家族;每周都可以从http://function.rcsb.org:8080/pdb/function_distribution/index.html.获得关于应该解决的最想要列表的更新信息人类基因组的测序为生物学家提供了理解生理过程和疾病状态的分子基础的新机会。为了充分利用这些机会,基因产物的三维结构需要提供适当的详细程度。由于蛋白质结构测定滞后于蛋白质序列测定,一个重要且持续的问题成为:我们从实验结构中对人类蛋白质组有多大程度的覆盖,我们可以通过建模推断出什么?或者,扭转这个问题:我们需要确定什么样的结构(最想要的名单)来加深我们对人类状况的理解?本文通过整合现有的使用比较功能特征相关的数据资源来解决这些问题,即描述所有类型蛋白质的生化过程、分子功能和细胞位置的基因本体论,以及酶的酶委员会分类。遗传病状态通过在线孟德尔人遗传资源联系在一起。读者可以在http://function.rcsb.org:8080/pdb/function_distribution/index.html.上向资源提出自己的问题该资源应被证明对结构基因组学特别有用,因为它正在努力进行大规模的结构确定,目的是改善对蛋白质功能空间的理解。
The bias in protein structure and function space resulting from experimental limitations and targeting of particular functional classes of proteins by structural biologists has long been recognized, but never continuously quantified. Using the Enzyme Commission and the Gene Ontology classifications as a reference frame, and integrating structure data from the Protein Data Bank (PDB), target sequences from the structural genomics projects, structure homology derived from the SUPERFAMILY database, and genome annotations from Ensembl and NCBI, we provide a quantified view, both at the domain and whole-protein levels, of the current and projected coverage of protein structure and function space relative to the human genome. Protein structures currently provide at least one domain that covers 37% of the functional classes identified in the genome; whole structure coverage exists for 25% of the genome. If all the structural genomics targets were solved (twice the current number of structures in the PDB), it is estimated that structures of one domain would cover 69% of the functional classes identified and complete structure coverage would be 44%. Homology models from existing experimental structures extend the 37% coverage to 56% of the genome as single domains and 25% to 31% for complete structures. Coverage from homology models is not evenly distributed by protein family, reflecting differing degrees of sequence and structure divergence within families. While these data provide coverage, conversely, they also systematically highlight functional classes of proteins for which structures should be determined. Current key functional families without structure representation are highlighted here; updated information on the “most wanted list” that should be solved is available on a weekly basis from http://function.rcsb.org:8080/pdb/function_distribution/index.html. The sequencing of the human genome provides biologists with new opportunities to understand the molecular basis of physiological processes and disease states. To take full advantage of these opportunities, the three-dimensional structures of the gene products are needed to provide the appropriate level of detail. Since protein structure determination lags behind protein sequence determination, an important and ongoing question becomes: what degree of coverage of the human proteome do we have from experimental structures, and what can we infer by modeling? Or, turning the question around: what structures do we need to determine (the “most wanted list”) to further our understanding of the human condition? This paper addresses these questions through integration of existing data resources correlated using comparative functional features, namely the Gene Ontology, which describes biochemical process, molecular function, and cellular location for all types of proteins, and the Enzyme Commission classification for enzymes. Genetic disease states are linked through the Online Mendelian Inheritance in Man resource. Readers can ask their own questions of the resource at http://function.rcsb.org:8080/pdb/function_distribution/index.html. The resource should prove particularly useful to the structural genomics community as it strives to undertake large-scale structure determination with a goal of improving the understanding of protein functional space.
DOI: 10.1093/bioinformatics/18.7.922
发表时间: 2002-07-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Liu, JF;Rost, B
通讯作者: Rost, B
DOI: 10.1093/nar/30.1.38
发表时间: 2002-01-01
影响因子: 14.9
作者:
Hubbard, T;Barker, D;Clamp, M
通讯作者: Clamp, M
DOI: 10.1093/bioinformatics/17.3.282
发表时间: 2001-03-01
期刊: BIOINFORMATICS
影响因子: 5.8
作者:
Li, WZ;Jaroszewski, L;Godzik, A
通讯作者: Godzik, A
DOI: 10.1093/nar/gkl929
发表时间: 2007-01-01
影响因子: 14.9
作者:
Bairoch, Amos;Bougueleret, Lydie;Zhang, Jian
通讯作者: Zhang, Jian
DOI: 10.1093/nar/gkp985
发表时间: 2010-01
影响因子: 14.9
作者:
Finn RD;Mistry J;Tate J;Coggill P;Heger A;Pollington JE;Gavin OL;Gunasekaran P;Ceric G;Forslund K;Holm L;Sonnhammer EL;Eddy SR;Bateman A
通讯作者: Bateman A