Comparative Analysis Of Completely Sequenced Genomes
Comparative Analysis Of Completely Sequenced Genomes
批准号:
9160910
负责人:
Eugene V Koonin
金额:
$30.47万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
AccountingAdoptedAffectAnimalsArchaeaBacteriaBacteriophagesBerylliumClustered Regularly Interspaced Short Palindromic RepeatsCollectionComputing MethodologiesDataDatabasesDouble Stranded DNA VirusDouble Stranded RNA VirusDouble-Stranded RNAElementsEukaryotaEvolutionFamilyGene FrequencyGenesGenetic VariationGenomeGenome engineeringGenomicsGoalsHorizontal Gene TransferImmune systemImmunoglobulin GenesIndividualLifeMethodsMinorityMobile Genetic ElementsModelingModificationNatureOrganismOrthologous GenePhylogenetic AnalysisPhylogenyPlant RootsPlasmidsProbabilityProcessProkaryotic CellsRNA VirusesRecording of previous eventsRelative (related person)ResearchRetroelementsRoleSamplingSingle Stranded DNA VirusSourceSurveysSystemTechnologyTectivirusesTheoretical StudiesTreesV(D)J RecombinationVertebratesViralVirionVirusadaptive immunitycomparativeds-DNAexperiencefusion genegenetic elementgenome editinggenome sequencinggenome-wideinsightlong term memorymathematical methodsmathematical modelparalogous geneplasmid DNArecombinasetooltrendvirome
中文摘要
细菌、古生物、真核生物和病毒的完全和几乎完全测序的基因组数据库迅速增长(已有数千个基因组,还有更多基因组正在进行中),为基因组研究创造了新的机遇和新的挑战。在过去的一年里,我们进行了各种研究,利用基因组信息来建立基因组进化的基本原则。
对感染真核宿主的病毒及其相关的移动元件进行了全面的系统基因组学分析。就物质丰度和遗传多样性而言,病毒和其他自私的遗传要素是生物圈中的主要实体。各种自私的元素寄生在所有细胞生命形式上。不同种类的病毒在原核生物和真核生物中的相对丰度有很大的差异。在原核生物中,绝大多数病毒具有双链(Ds)DNA基因组,相当数量的单链(Ss)DNA病毒和有限的RNA病毒存在。相反,在真核生物中,尽管ssDNA和dsDNA病毒也很常见,但RNA病毒占病毒多样性的大部分。系统基因组学分析为真核病毒的主要类别的起源提供了切实的线索,特别是它们可能起源于原核生物。具体地说,真核生物的正链RNA病毒的祖先基因组可能是从原核反转录元件和细菌的基因重新组装而来的,尽管不能排除这类病毒的原始起源。不同组的双链RNA病毒要么来自dsRNA噬菌体,要么来自正链RNA病毒。真核单链DNA病毒显然是通过原核滚动环复制质粒和正链RNA病毒的基因融合进化而来的。真核dsDNA病毒的不同家族似乎至少在两次独立的情况下起源于特定的噬菌体群。Polintons是已知的最大的真核转座子,预计也会形成病毒颗粒,很可能是细菌顶盖病毒和几组真核dsDNA病毒之间的进化中间体,包括拟议的“Megavirales”目,它联合了不同的大型和巨型病毒家族。引人注目的是,所有类别的真核病毒的进化似乎涉及到来自不同来源的结构和复制基因模块之间的融合,以及对不同基因的额外获取。
我们开发了适应性免疫系统进化的一般场景,并可能从移动遗传元素发展出其他基因组操作机制。原核生物和动物中的适应性免疫系统通过改变特定的基因组位点来产生长期记忆,例如通过将外源DNA片段插入原核生物中成簇的规则间隔短回文重复序列(CRISPR)中,以及通过脊椎动物中免疫球蛋白基因的V(D)J重组。值得注意的是,来自不相关的可移动遗传元件的重组酶在原核生物和脊椎动物的适应性免疫系统中都具有重要的作用。在细胞生命形式中普遍存在的移动元件,为基因组工程提供了唯一已知的、自然进化的工具,并成功地被先天性免疫系统和基因组编辑技术采用。
我们对细菌和古生菌超基因组的进化进行了理论研究。由于原核生物基因组经历了快速的基因流动,选择可能在比单个基因组更高的水平上发挥作用。我们探索了一种分布式基因组的定量模型,在该模型中,基因组组通过从我们称为超基因组的固定储存库中获取基因来进化。以前对超基因组性质的理解将基因组视为随机的、独立的基因集合,并假设超基因组由少量同质亚库组成。在这里,我们将探讨放松这两个假设的后果。
我们考察了几种估算超基因组大小和组成的方法。这些方法假设基因组要么是超基因组的随机、独立样本,要么是通过从储藏库随机抽样从已知树上的共同祖先进化而来的。该储集层被假定为一组均质子储集层,或者由具有伽玛分布的增益概率的基因组成。经验基因频率被用来直接或首先计算数据的可能性,以重建基因增益的历史,然后计算重建的增益数目的可能性。直接使用经验基因频率估计超基因组大小对于模型的选择并不稳健。相比之下,使用基因频率和系统发育树重建多个基因增益产生了对超基因组大小的可靠估计,并表明同源超基因组比具有伽玛分布增益概率的超基因组更符合数据。
综上所述,这些研究促进了对不同生命形式,特别是病毒和移动元素中基因组进化的现有理解,并为基因组进化的一般原理提供了新的见解。
英文摘要
The rapidly growing database of completely and nearly completely sequenced genomes of bacteria, archaea, eukaryotes and viruses (several thousand genomes already available and many more in progress) creates both new opportunities and new challenges for genome research. During the last year, we performed a variety of studies that took advantage of the genomic information to establish fundamental principles of genome evolution.
A comprehensive phylogenomic analysis of viruses infecting eukaryotic hosts and the related mobile elements was performed. Viruses and other selfish genetic elements are dominant entities in the biosphere, with respect to both physical abundance and genetic diversity. Various selfish elements parasitize on all cellular life forms. The relative abundances of different classes of viruses are dramatically different between prokaryotes and eukaryotes. In prokaryotes, the great majority of viruses possess double-stranded (ds) DNA genomes, with a substantial minority of single-stranded (ss) DNA viruses and only limited presence of RNA viruses. In contrast, in eukaryotes, RNA viruses account for the majority of the virome diversity although ssDNA and dsDNA viruses are common as well. Phylogenomic analysis yields tangible clues for the origins of major classes of eukaryotic viruses and in particular their likely roots in prokaryotes. Specifically, the ancestral genome of positive-strand RNA viruses of eukaryotes might have been assembled de novo from genes derived from prokaryotic retroelements and bacteria although a primordial origin of this class of viruses cannot be ruled out. Different groups of double-stranded RNA viruses derive either from dsRNA bacteriophages or from positive-strand RNA viruses. The eukaryotic ssDNA viruses apparently evolved via a fusion of genes from prokaryotic rolling circle-replicating plasmids and positive-strand RNA viruses. Different families of eukaryotic dsDNA viruses appear to have originated from specific groups of bacteriophages on at least two independent occasions. Polintons, the largest known eukaryotic transposons, predicted to also form virus particles, most likely, were the evolutionary intermediates between bacterial tectiviruses and several groups of eukaryotic dsDNA viruses including the proposed order "Megavirales" that unites diverse families of large and giant viruses. Strikingly, evolution of all classes of eukaryotic viruses appears to have involved fusion between structural and replicative gene modules derived from different sources along with additional acquisitions of diverse genes.
We developed a general scenario of evolution of adaptive immunity systems and possibly other genome manipulation machineries from mobile genetic elements. Adaptive immune systems in prokaryotes and animals give rise to long-term memory through modification of specific genomic loci, such as by insertion of foreign (viral or plasmid) DNA fragments into clustered regularly interspaced short palindromic repeat (CRISPR) loci in prokaryotes and by V(D)J recombination of immunoglobulin genes in vertebrates. Strikingly, recombinases derived from unrelated mobile genetic elements have essential roles in both prokaryotic and vertebrate adaptive immune systems. Mobile elements, which are ubiquitous in cellular life forms, provide the only known, naturally evolved tools for genome engineering that are successfully adopted by both innate immune systems and genome-editing technologies.
We performed a theoretical study of the evolution of bacterial and archaeal supergenomes. Because prokaryotic genomes experience a rapid flux of genes, selection may act at a higher level than an individual genome. We explore a quantitative model of the distributed genome whereby groups of genomes evolve by acquiring genes from a fixed reservoir which we denote as supergenome. Previous attempts to understand the nature of the supergenome treated genomes as random, independent collections of genes and assumed that the supergenome consists of a small number of homogeneous sub-reservoirs. Here we explore the consequences of relaxing both assumptions.
We surveyed several methods for estimating the size and composition of the supergenome. The methods assumed that genomes were either random, independent samples of the supergenome or that they evolved from a common ancestor along a known tree via stochastic sampling from the reservoir. The reservoir was assumed to be either a collection of homogeneous sub-reservoirs or alternatively composed of genes with Gamma distributed gain probabilities. Empirical gene frequencies were used to either compute the likelihood of the data directly or first to reconstruct the history of gene gains and then compute the likelihood of the reconstructed numbers of gains. Supergenome size estimates using the empirical gene frequencies directly are not robust with respect to the choice of the model. By contrast, using the gene frequencies and the phylogenetic tree to reconstruct multiple gene gains produces reliable estimates of the supergenome size and indicates that a homogeneous supergenome is more consistent with the data than a supergenome with Gamma distributed gain probabilities.
Taken together, these studies advance the existing understanding of the genome evolution in diverse life forms, in particular viruses and mobile elements, and provide new insights into general principles of genome evolution.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Finding Protein Sequence Motifs--methods And Application
-
批准号:6681337
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--Methods and Application
-
批准号:6988455
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:7969213
-
项目类别:
-
资助金额:$195.34万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--methods And Applications
-
批准号:8943217
-
项目类别:
-
资助金额:$30.99万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--methods And Applications
-
批准号:7735068
-
项目类别:
-
资助金额:$32.76万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--methods And Applications
-
批准号:7594460
-
项目类别:
-
资助金额:$31.78万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--methods And Applications
-
批准号:9555730
-
项目类别:
-
资助金额:$31.91万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
COMPARATIVE ANALYSIS OF COMPLETELY SEQUENCED GENOMES
-
批准号:6111075
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:6988458
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:7316251
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
COMPARATIVE ANALYSIS OF COMPLETELY SEQUENCED GENOMES
-
批准号:6432755
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
COMPARATIVE ANALYSIS OF COMPLETELY SEQUENCED GENOMES
-
批准号:6554459
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--methods And Applications
-
批准号:8344941
-
项目类别:
-
资助金额:$119.92万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:8943219
-
项目类别:
-
资助金额:$299.48万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:9362440
-
项目类别:
-
资助金额:$274.39万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--Methods And Applications
-
批准号:10691115
-
项目类别:
-
资助金额:$41.53万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--Methods And Applications
-
批准号:10925004
-
项目类别:
-
资助金额:$44.27万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:10927035
-
项目类别:
-
资助金额:$351.37万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:8149599
-
项目类别:
-
资助金额:$215.46万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
FINDING PROTEIN SEQUENCE MOTIFS--METHODS AND APPLICATIONS
-
批准号:6290486
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
海外基金