Comparative Analysis Of Completely Sequenced Genomes
Comparative Analysis Of Completely Sequenced Genomes
批准号:
7316251
负责人:
Eugene V Koonin
金额:
$0.0万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
中文摘要
点击翻译按钮获取中文摘要
英文摘要
The rapidly growing database of completely sequenced genomes of bacteria, archaea and eukaryotes (over 200 genomes available by the end of 2004 and many more in progress) creates both new opportunities and new challenges for genome research. During the last year, we performed several studies that took advantage of the genomic information to establish fundamental principles of genome evolution and function. By comparing sequences of human, mouse and rat orthologous genes, we show that in 5'-untranslated regions (5'-UTRs) of mammalian cDNAs but not in 3'-UTRs or coding sequences, AUG is conserved to a significantly greater extent than any of the other 63 nt triplets. Qualitatively similar results were obtained by comparison of orthologous genes from different species of the yeast genus Saccharomyces. Together with the observation that mammalian and yeast 5'-UTRs are significantly depleted in overall AUG content, these findings suggest that AUG triplets in 5'-UTRs are subject to the pressure of purifying selection in two opposite directions: the uAUGs that have no specific function tend to be deleterious and get eliminated during evolution, whereas those uAUGs that do serve a function are conserved. Most probably, the principal role of the conserved uAUGs is attenuation of translation at the initiation stage, which is often additionally regulated by alternative splicing in the mammalian 5'-UTRs. In another project, we assessed the extent of ancestral paralogy, which dates back to the last common ancestor of all eukaryotes, and examine the origins of the ancestral paralogs and their potential roles in the emergence of the eukaryotic cell complexity. A parsimonious reconstruction of ancestral gene repertoires shows that 4137 orthologous gene sets in the last eukaryotic common ancestor (LECA) map back to 2150 orthologous sets in the hypothetical first eukaryotic common ancestor (FECA) [paralogy quotient (PQ) of 1.92]. Analogous reconstructions show significantly lower levels of paralogy in prokaryotes, 1.19 for archaea and 1.25 for bacteria. The only functional class of eukaryotic proteins with a significant excess of paralogous clusters over the mean includes molecular chaperones and proteins with related functions. Almost all genes in this category underwent multiple duplications during early eukaryotic evolution. In structural terms, the most prominent sets of paralogs are superstructure-forming proteins with repetitive domains, such as WD-40 and TPR. In addition to the ancestral paralogs which evolved via duplication at the onset of eukaryotic evolution, numerous pseudoparalogs were detected, i.e. homologous genes that apparently were acquired by early eukaryotes via different routes, including horizontal gene transfer (HGT) from diverse bacteria. The results of this study demonstrate a major increase in the level of gene paralogy as a hallmark of the early evolution of eukaryotes. We also studied universal trends in the evolution of amino acid composition of proteins. We compared sets of orthologous proteins encoded by triplets of closely related genomes from 15 taxa representing all three domains of life (Bacteria, Archaea and Eukaryota), and used phylogenies to polarize amino acid substitutions. Cys, Met, His, Ser and Phe accrue in at least 14 taxa, whereas Pro, Ala, Glu and Gly are consistently lost. The same nine amino acids are currently accrued or lost in human proteins, as shown by analysis of non-synonymous single-nucleotide polymorphisms. All amino acids with declining frequencies are thought to be among the first incorporated into the genetic code; conversely, all amino acids with increasing frequencies, except Ser, were probably recruited late. Thus, expansion of initially under-represented amino acids, which began over 3,400 million years ago, apparently continues to this day.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
Finding Protein Sequence Motifs--methods And Application
-
批准号:6681337
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--Methods and Application
-
批准号:6988455
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:7969213
-
项目类别:
-
资助金额:$195.34万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--methods And Applications
-
批准号:8943217
-
项目类别:
-
资助金额:$30.99万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:9160910
-
项目类别:
-
资助金额:$30.47万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--methods And Applications
-
批准号:7735068
-
项目类别:
-
资助金额:$32.76万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--methods And Applications
-
批准号:7594460
-
项目类别:
-
资助金额:$31.78万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--methods And Applications
-
批准号:9555730
-
项目类别:
-
资助金额:$31.91万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
COMPARATIVE ANALYSIS OF COMPLETELY SEQUENCED GENOMES
-
批准号:6111075
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:6988458
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
COMPARATIVE ANALYSIS OF COMPLETELY SEQUENCED GENOMES
-
批准号:6432755
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
COMPARATIVE ANALYSIS OF COMPLETELY SEQUENCED GENOMES
-
批准号:6554459
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:9362440
-
项目类别:
-
资助金额:$274.39万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--methods And Applications
-
批准号:8344941
-
项目类别:
-
资助金额:$119.92万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:8943219
-
项目类别:
-
资助金额:$299.48万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--Methods And Applications
-
批准号:10691115
-
项目类别:
-
资助金额:$41.53万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:10927035
-
项目类别:
-
资助金额:$351.37万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--Methods And Applications
-
批准号:10925004
-
项目类别:
-
资助金额:$44.27万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:8149599
-
项目类别:
-
资助金额:$215.46万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
FINDING PROTEIN SEQUENCE MOTIFS--METHODS AND APPLICATIONS
-
批准号:6290486
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
国内基金
海外基金
登录
查看更多内容
Scalable Learning and Optimization: High-dimensional Models and Online Decision-Making Strategies for Big Data Analysis
-
批准号:--
-
项目类别:合作创新研究团队
-
资助金额:--
-
批准年份:2024
-
负责人:姚韬
-
依托单位:
Intelligent Patent Analysis for Optimized Technology Stack Selection:Blockchain BusinessRegistry Case Demonstration
-
批准号:--
-
项目类别:外国学者研究基金项目
-
资助金额:--
-
批准年份:2024
-
负责人:USHARANI HAREESH GOVINDARA JAN
-
依托单位:
基于Meta-analysis的新疆棉花灌水增产模型研究
-
批准号:41601604
-
项目类别:青年科学基金项目
-
资助金额:22.0万元
-
批准年份:2016
-
负责人:赵爱琴
-
依托单位:
大规模微阵列数据组的meta-analysis方法研究
-
批准号:31100958
-
项目类别:青年科学基金项目
-
资助金额:20.0万元
-
批准年份:2011
-
负责人:赵洪雅
-
依托单位:
用“后合成核磁共振分析”(retrobiosynthetic NMR analysis)技术阐明青蒿素生物合成途径
-
批准号:30470153
-
项目类别:面上项目
-
资助金额:22.0万元
-
批准年份:2004
-
负责人:刘本叶
-
依托单位: