Finding Protein Sequence Motifs--Methods And Applications
Finding Protein Sequence Motifs--Methods And Applications
批准号:
10925004
负责人:
Eugene V Koonin
金额:
$44.27万
依托单位国家:
美国
项目类别:
财政年份:
--
资助国家:
美国
项目状态:
未结题
起止时间:
至
关键词:
Amino Acid MotifsAmino Acid SequenceArchaeal VirusesArchitectureArtificial IntelligenceBacteriaBacteriophagesBindingCRISPR-associated transposonsCase StudyCatalytic DomainClustered Regularly Interspaced Short Palindromic RepeatsCollaborationsCollectionComplementComputing MethodologiesCustomDNA LigasesDNA Transposable ElementsDNA-Directed DNA PolymeraseDNA-Directed RNA PolymeraseDatabasesDetectionDevelopmentDissectionEnzymesEssential GenesEukaryotaEvolutionFamilyGenesGenomeGenomicsGoalsGuide RNAHomologous GeneHomology ModelingHost DefenseInvestigationLaboratoriesLibrariesLifeMajor Core ProteinMethodsModelingNucleotidesOrganismOrthopoxvirusPatternPhosphotransferasesPositioning AttributePoxviridaeProcollagen-Proline DioxygenaseProkaryotic CellsProtein AnalysisProteinsProteomeReproductionResolutionRoleSOS ResponseSequence AnalysisSiteSourceStructural ModelsStructureSurveysSystemTertiary Protein StructureToxinTyrosineViral GenomeViral ProteinsVirionVirusWorkadaptive immunityantitoxinarms racecomputational pipelinesdatabase structuredeep learningds-DNAhuman pathogenmarkov modelmolecular sequence databasenovelprofessorprotein data bankprotein foldingprotein structureprotein structure predictionrecombinaserecruittrendubiquitin isopeptidase
中文摘要
基因组序列和蛋白质结构在过去十年中的快速积累已经被序列数据库搜索方法以及蛋白质结构预测的重大进展所证实。NCBI开发的强大的位置特异性迭代BLAST(PSI-BLAST)方法构成了我们蛋白质基序分析工作的基础。此外,隐马尔可夫模型(HMM)、在HHSearch方法中实现的蛋白质图谱对图谱比较、蛋白质结构比较方法、蛋白质结构的同源性建模和基因组背景分析被广泛且越来越多地应用。此外,蛋白质结构域谱的定制库以及用于新结构域识别的计算管道已经开发和应用。最近,这些用于蛋白质基序搜索的方法正在被深度学习计算方法所补充,特别是AlphaFold 2,这是一种强大的蛋白质结构建模方法。
在回顾的一年中,我们继续研究蛋白质结构域,特别是那些在原核生物和真核生物病毒基因组中编码的蛋白质结构域,以及涉及细菌防御病毒的结构域。这些研究的范围通过AlphaFold 2的广泛使用而大大扩展。
作为AlphaFold 2应用于预测快速进化病毒蛋白的结构和功能的案例研究,我们对正痘病毒的蛋白质组进行了建模,正痘病毒是一组重要的病毒,包括主要的人类病原体。具有大的双链DNA基因组的病毒在进化的不同阶段从宿主那里捕获了大部分基因。许多病毒基因的起源很容易通过与细胞同源物的显著序列相似性来检测。特别是,对于病毒酶,如DNA和RNA聚合酶或核苷酸激酶,情况就是这样,它们在被祖先病毒捕获后保留了它们的催化活性。然而,很大一部分病毒基因没有容易检测到的细胞同源物,这意味着它们的起源仍然是个谜。我们探索了编码在正痘病毒基因组中的这种蛋白质的潜在来源,正痘病毒是一种经过彻底研究的病毒属,包括主要的人类病原体。为此,我们使用AlphaFold 2来预测正痘病毒编码的所有214种蛋白质的结构。在未知来源的蛋白质中,结构预测为其中14个蛋白质提供了明确的来源指示,并验证了先前通过序列分析得出的几个推论。一个值得注意的新趋势是,来自细胞生物体的酶在病毒繁殖中的非酶结构作用,伴随着催化位点的破坏和整体的急剧分化,从而排除了在序列水平上的同源性检测。在16种被发现是失活酶衍生物的正痘病毒蛋白中,有痘病毒复制持续合成因子A20,它是一种失活的NAD依赖性DNA连接酶;主要核心蛋白A3,它是一种失活的去泛素酶; F11,它是一种失活的脯氨酰羟化酶;以及更多类似的情况。对于近三分之一的正痘病毒病毒体蛋白,没有发现显着相似的结构,这表明随后的主要结构重排,产生独特的蛋白质折叠的exaptation。
在另一项研究中,我们着手确定由细菌和古细菌病毒编码的抗CRISPR蛋白(ACR)的起源。大多数Acr是小的非酶蛋白,其通过与Cas效应蛋白结合来消除CRISPR活性。Acr进化得很快,这是由于与各自的CRISPR-Cas系统的军备竞赛,这阻碍了通过序列比较阐明它们的进化起源。我们使用AlphaFold 2对3693个实验表征和预测的Acr进行了全面的结构建模,然后与蛋白质数据库中的蛋白质结构进行了比较。通过序列相似性对Acr进行聚类分析,得到了363个高质量的结构模型,共包含102个Acr家族。结构比较允许确定同源的13个这些家庭可能是祖先的Acr。尽管有限的程度上的结构保守,推断的起源的Acr显示出不同的趋势,特别是,招聘的毒素和抗毒素和SOS修复系统组件的Acr功能。
本研究与麻省理工学院布罗德研究所和哈佛大学张峰教授的实验室合作,对Tn 7类转座子的靶选择蛋白的结构和功能进行了研究。为了传播,转座子必须在不破坏必需基因的情况下整合到靶位点,同时避开宿主防御系统。Tn 7样转座子采用多种机制进行靶位点选择,包括蛋白质引导的靶向,以及CRISPR相关转座子(CAST)中的RNA引导的靶向。结合基因组学和结构分析,我们进行了广泛的调查目标选择,揭示了不同的机制Tn 7识别靶位点,包括以前未表征的目标选择蛋白中发现的新发现的转座因子(TE)。我们的实验特征在于CAST I-D系统和Tn 6022样转座子,使用TnSF,其中包含一个失活的酪氨酸重组酶结构域,靶向comM基因。此外,我们确定了一个非Tn 7转座子,Tsy,编码一个同源的TnSF与一个活跃的酪氨酸重组酶结构域,我们也显示插入comM。我们的研究结果表明,Tn 7转座子采用模块化的架构和增选目标选择器从各种来源,以优化目标选择和驱动转座子传播。
英文摘要
The rapid accumulation of genome sequences and protein structures during the last decade has been paralleled by major advances in sequence database search methods as well as protein structure prediction. The powerful Position-Specific Iterating BLAST (PSI-BLAST) method developed at the NCBI forms the basis of our work on protein motif analysis. In addition, Hidden Markov Models (HMM), protein profile-against-profile comparison implemented in the HHSearch method, protein structure comparison methods, homology modeling of protein structure and genome context analysis were extensively and increasingly applied. Furthermore, custom libraries of protein domain profiles as well as computational pipelines for novel domain identification have been developed and applied. Lately, these methods for protein motif search are being complemented by deep learning computational methods, in particular, AlphaFold2, a powerful method for protein structure modeling.
During the year under review, we have continued our investigation of the proteins domains, particularly, those that are encoded in the genomes of viruses of prokaryotes and eukaryotes as well as domains involved in the defense of bacteria against viruses. The scope of these studies was substantially expanded through extensive use of AlphaFold2.
As a case study for the application of AlphaFold2 to predict structures and functions of fast-evolving virus proteins, we modeled the proteome of orthopoxviruses, an important group of viruses that includes major human pathogens. Viruses with large, double-stranded DNA genomes captured the majority of their genes from their hosts at different stages of evolution. The origins of many virus genes are readily detected through significant sequence similarity with cellular homologs. In particular, this is the case for virus enzymes, such as DNA and RNA polymerases or nucleotide kinases, that retain their catalytic activity after capture by an ancestral virus. However, a large fraction of virus genes have no readily detectable cellular homologs, meaning that their origins remain enigmatic. We explored the potential origins of such proteins that are encoded in the genomes of orthopoxviruses, a thoroughly studied virus genus that includes major human pathogens. To this end, we used AlphaFold2 to predict the structures of all 214 proteins that are encoded by orthopoxviruses. Among the proteins of unknown provenance, structure prediction yielded clear indications of origin for 14 of them and validated several inferences that were previously made via sequence analysis. A notable emerging trend is the exaptation of enzymes from cellular organisms for nonenzymatic, structural roles in virus reproduction that is accompanied by the disruption of catalytic sites and by an overall drastic divergence that precludes homology detection at the sequence level. Among the 16 orthopoxvirus proteins that were found to be inactivated enzyme derivatives are the poxvirus replication processivity factor A20, which is an inactivated NAD-dependent DNA ligase; the major core protein A3, which is an inactivated deubiquitinase; F11, which is an inactivated prolyl hydroxylase; and more similar cases. For nearly one-third of the orthopoxvirus virion proteins, no significantly similar structures were identified, suggesting exaptation with subsequent major structural rearrangement that yielded unique protein folds.
In another study, we set out to identify the origins of anti-CRISPR proteins (ACRs) encoded by bacterial and archaeal viruses. The majority of the Acrs are small, non-enzymatic proteins that abrogate CRISPR activity by binding to Cas effector proteins. The Acrs evolve fast, due to the arms race with the respective CRISPR-Cas systems, which hampers the elucidation of their evolutionary origins by sequence comparison. We performed comprehensive structural modeling using AlphaFold2 for 3693 experimentally characterized and predicted Acrs, followed by a comparison to the protein structures in the Protein Data Bank database. After clustering the Acrs by sequence similarity, 363 high-quality structural models were obtained that accounted for 102 Acr families. Structure comparisons allowed the identification of homologs for 13 of these families that could be ancestors of the Acrs. Despite the limited extent of structural conservation, the inferred origins of Acrs show distinct trends, in particular, recruitment of toxins and antitoxins and SOS repair system components for the Acr function.
In collaboration with the laboratory of Professor Feng Zhang at the Broad Insitute of MIT and Harvard, we investigated the structures and functional of target selector proteins of Tn7-like transposons. To spread, transposons must integrate into target sites without disruption of essential genes while avoiding host defense systems. Tn7-like transposons employ multiple mechanisms for target-site selection, including protein-guided targeting and, in CRISPR-associated transposons (CASTs), RNA-guided targeting. Combining phylogenomic and structural analyses, we conducted a broad survey of target selectors, revealing diverse mechanisms used by Tn7 to recognize target sites, including previously uncharacterized target-selector proteins found in newly discovered transposable elements (TEs). We experimentally characterized a CAST I-D system and a Tn6022-like transposon that uses TnsF, which contains an inactivated tyrosine recombinase domain, to target the comM gene. Additionally, we identified a non-Tn7 transposon, Tsy, encoding a homolog of TnsF with an active tyrosine recombinase domain, which we show also inserts into comM. Our findings show that Tn7 transposons employ modular architecture and co-opt target selectors from various sources to optimize target selection and drive transposon spread.
期刊论文(32)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1186/1745-6150-8-15
发表时间:
2013-06-15
期刊:
Biology direct
影响因子:
5.5
作者:
[Anantharaman V, Makarova KS, Burroughs AM, Koonin EV, Aravind L]
通讯作者:
Aravind L
DOI:
10.1016/j.tim.2015.05.005
发表时间:
2015-09
期刊:
Trends in microbiology
影响因子:
15.9
作者:
[Römling U, Galperin MY]
通讯作者:
Galperin MY
DOI:
10.1186/s13062-017-0185-2
发表时间:
2017-05-25
期刊:
Biology direct
影响因子:
5.5
作者:
[Mekhedov SL, Makarova KS, Koonin EV]
通讯作者:
Koonin EV
The CMG (CDC45/RecJ, MCM, GINS) complex is a conserved component of the DNA replication system in all archaea and eukaryotes.
CMG(CDC45/RECJ,MCM,GINS)复合物是所有古细菌和真核生物中DNA复制系统的保守组成部分。
DOI:
10.1186/1745-6150-7-7
发表时间:
2012-02-13
期刊:
Biology direct
影响因子:
5.5
作者:
[Makarova KS, Koonin EV, Kelman Z]
通讯作者:
Kelman Z
DOI:
10.1186/s12915-020-00885-2
发表时间:
2020-11-04
期刊:
BMC biology
影响因子:
5.4
作者:
[Bell RT, Wolf YI, Koonin EV]
通讯作者:
Koonin EV
共 17 条
Finding Protein Sequence Motifs--Methods and Application
-
批准号:6988455
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--methods And Application
-
批准号:6681337
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:7969213
-
项目类别:
-
资助金额:$195.34万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--methods And Applications
-
批准号:8943217
-
项目类别:
-
资助金额:$30.99万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:9160910
-
项目类别:
-
资助金额:$30.47万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--methods And Applications
-
批准号:9555730
-
项目类别:
-
资助金额:$31.91万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--methods And Applications
-
批准号:7594460
-
项目类别:
-
资助金额:$31.78万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--methods And Applications
-
批准号:7735068
-
项目类别:
-
资助金额:$32.76万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
COMPARATIVE ANALYSIS OF COMPLETELY SEQUENCED GENOMES
-
批准号:6111075
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:6988458
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:7316251
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
COMPARATIVE ANALYSIS OF COMPLETELY SEQUENCED GENOMES
-
批准号:6432755
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
COMPARATIVE ANALYSIS OF COMPLETELY SEQUENCED GENOMES
-
批准号:6554459
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--methods And Applications
-
批准号:8344941
-
项目类别:
-
资助金额:$119.92万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:8943219
-
项目类别:
-
资助金额:$299.48万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:9362440
-
项目类别:
-
资助金额:$274.39万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Finding Protein Sequence Motifs--Methods And Applications
-
批准号:10691115
-
项目类别:
-
资助金额:$41.53万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:10927035
-
项目类别:
-
资助金额:$351.37万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
Comparative Analysis Of Completely Sequenced Genomes
-
批准号:8149599
-
项目类别:
-
资助金额:$215.46万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
FINDING PROTEIN SEQUENCE MOTIFS--METHODS AND APPLICATIONS
-
批准号:6290486
-
项目类别:
-
资助金额:$0.0万
-
财政年份:--
-
负责人:Eugene V Koonin
-
依托单位:
海外基金