The Sorcerer II Global Ocean Sampling expedition: expanding the universe of protein families.

The Sorcerer II Global Ocean Sampling expedition: expanding the universe of protein families.
复制标题

DOI:
10.1371/journal.pbio.0050016
复制
发表时间:
2007-03
期刊:
影响因子:
9.8
通讯作者:
Venter JC
Venter JC
中科院分区:
生物学1区
文献类型:
--
作者:
Yooseph S;Sutton G;Rusch DB;Halpern AL;Williamson SJ;Remington K;Eisen JA;Heidelberg KB;Manning G;Li W;Jaroszewski L;Cieplak P;Miller CS;Li H;Mashiyama ST;Joachimiak MP;van Belle C;Chandonia JM;Soergel DA;Zhai Y;Natarajan K;Lee S;Raphael BJ;Bafna V;Friedman R;Brenner SE;Godzik A;Eisenberg D;Dixon JE;Taylor SS;Strausberg RL;Frazier M;Venter JC

文献摘要

参考文献

被引文献

相似文献

基于微生物种群shot弹枪测序的宏基因组学项目会洞悉蛋白质家族。采样(GOS)序列。仅确定了由GOS序列组成的大型群集,其中1,700个没有可检测到的仅GOS家族的同源物。现在。在其他王国中,有大约6,000个序列(ORFAN)在迄今与已知蛋白质相似的文献中,GOS数据集的匹配也几乎可以改善远程同源性检测。当前的蛋白质数量,预测的GOS蛋白还为已知的蛋白质家族增添了很多多样性,并阐明了这些观察结果。使用几个蛋白质家族,包括磷酸酶,蛋白质,紫外线 - 辐照DNA损伤修复酶,谷氨酰胺合成酶和RUBISCO,GOS数据添加的多样性对选择实验结构的靶标有含义。正在以线性或几乎是线性的添加新序列发现新家庭,这意味着我们仍然是远没有发现自然界中的所有蛋白质家族。 宏基因组学的迅速新兴领域旨在检查组织社区的基因组内容,以了解其在生态系统中的角色和相互作用,鉴于微生物在许多生态系统中的广泛角色,对微生物群落的宏观学研究将揭示蛋白质家族和蛋白质家族和蛋白质家庭和洞察力它们的进化。多样性。一种这样的技术 - 刺激性的测序,使DNA序列随机取样,以检查微生物群落中存在的基因组材料。我们的分析预测了GOS数据中有超过600万个蛋白质,这些蛋白质的数量是当前数据库中存在的蛋白质数量的两倍。已知的蛋白质家族和几乎所有已知的蛋白质家族都与当前已知的蛋白质没有相似之处,因此预计这些新家族也比预期的发现以前被认为是王国的几个蛋白质领域在其他王国中有GOS的例子。蛋白质家族分析,并表明我们距离对自然界中存在的所有蛋白质家族进行取样还有很长的路要走。 GOS数据确定了612万个蛋白质涵盖了几乎所有已知的原核生物蛋白家族,这几乎是已知蛋白质的数量。
Metagenomics projects based on shotgun sequencing of populations of micro-organisms yield insight into protein families. We used sequence similarity clustering to explore proteins with a comprehensive dataset consisting of sequences from available databases together with 6.12 million proteins predicted from an assembly of 7.7 million Global Ocean Sampling (GOS) sequences. The GOS dataset covers nearly all known prokaryotic protein families. A total of 3,995 medium- and large-sized clusters consisting of only GOS sequences are identified, out of which 1,700 have no detectable homology to known families. The GOS-only clusters contain a higher than expected proportion of sequences of viral origin, thus reflecting a poor sampling of viral diversity until now. Protein domain distributions in the GOS dataset and current protein databases show distinct biases. Several protein domains that were previously categorized as kingdom specific are shown to have GOS examples in other kingdoms. About 6,000 sequences (ORFans) from the literature that heretofore lacked similarity to known proteins have matches in the GOS data. The GOS dataset is also used to improve remote homology detection. Overall, besides nearly doubling the number of current proteins, the predicted GOS proteins also add a great deal of diversity to known protein families and shed light on their evolution. These observations are illustrated using several protein families, including phosphatases, proteases, ultraviolet-irradiation DNA damage repair enzymes, glutamine synthetase, and RuBisCO. The diversity added by GOS data has implications for choosing targets for experimental structure characterization as part of structural genomics efforts. Our analysis indicates that new families are being discovered at a rate that is linear or almost linear with the addition of new sequences, implying that we are still far from discovering all protein families in nature. The rapidly emerging field of metagenomics seeks to examine the genomic content of communities of organisms to understand their roles and interactions in an ecosystem. Given the wide-ranging roles microbes play in many ecosystems, metagenomics studies of microbial communities will reveal insights into protein families and their evolution. Because most microbes will not grow in the laboratory using current cultivation techniques, scientists have turned to cultivation-independent techniques to study microbial diversity. One such technique—shotgun sequencing—allows random sampling of DNA sequences to examine the genomic material present in a microbial community. We used shotgun sequencing to examine microbial communities in water samples collected by the Sorcerer II Global Ocean Sampling (GOS) expedition. Our analysis predicted more than six million proteins in the GOS data—nearly twice the number of proteins present in current databases. These predictions add tremendous diversity to known protein families and cover nearly all known prokaryotic protein families. Some of the predicted proteins had no similarity to any currently known proteins and therefore represent new families. A higher than expected fraction of these novel families is predicted to be of viral origin. We also found that several protein domains that were previously thought to be kingdom specific have GOS examples in other kingdoms. Our analysis opens the door for a multitude of follow-up protein family analyses and indicates that we are a long way from sampling all the protein families that exist in nature. The GOS data identified 6.12 million predicted proteins covering nearly all known prokaryotic protein families, and several new families. This almost doubles the number of known proteins and shows that we are far from identifying all the proteins in nature.
DOI: 10.1046/j.1365-2958.2003.03657.x
发表时间: 2003-09-01
影响因子: 3.6
作者:
Boitel, B;Ortiz-Lombardía, M;Alzari, PM
通讯作者: Alzari, PM
DOI: 10.1126/science.286.5439.509
发表时间: 1999-10-15
期刊: SCIENCE
影响因子: 56.9
作者:
Barabási, AL;Albert, R
通讯作者: Albert, R
DOI: 10.1093/nar/gki034
发表时间: 2005-01-01
影响因子: 14.9
作者:
Bru C;Courcelle E;Carrère S;Beausse Y;Dalmar S;Kahn D
通讯作者: Kahn D
DOI: 10.1093/nar/gkp985
发表时间: 2010-01
影响因子: 14.9
作者:
Finn RD;Mistry J;Tate J;Coggill P;Heger A;Pollington JE;Gavin OL;Gunasekaran P;Ceric G;Forslund K;Holm L;Sonnhammer EL;Eddy SR;Bateman A
通讯作者: Bateman A
DOI: 10.1186/gb-2004-5-5-r35
发表时间: 2004
期刊: Genome biology
影响因子: 12.3
作者:
Bowers PM;Pellegrini M;Thompson MJ;Fierro J;Yeates TO;Eisenberg D
通讯作者: Eisenberg D