BBSRC-NSF/BIO: An AI-based domain classification platform for 200 million 3D-models of proteins to reveal protein evolution
BBSRC-NSF/BIO: An AI-based domain classification platform for 200 million 3D-models of proteins to reveal protein evolution
批准号:
BB/Y001117/1
负责人:
Christine Orengo
金额:
$34.21万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2024
资助国家:
英国
项目状态:
未结题
起止时间:
2024 至 --
中文摘要
蛋白质在生命中最重要的过程中起着重要的作用,如营养物质的消化、免疫反应和细胞调节。它们由长聚合物组成,折叠成紧密的球形,称为域。大多数蛋白质至少有两个结构域,有些由几十个结构域组成。领域往往与特定的功能相关联,尽管有时一个重要的功能将由多个领域组合而成。三维结构数据和模型对于检测与域功能相关的口袋和表面特征特别有价值。确定组成结构域的结构和方向对于理解蛋白质的整体功能以及与之相关的动态构象变化非常重要。直到最近,蛋白质的结构数据非常稀少,只有不到1%的已知蛋白质被实验表征。虽然当一个近亲属的结构已知时,结构可以有合理的准确性预测,但对于很大比例的蛋白质,这样的数据不存在。即使是像人类或小麦这样重要的生物,也只有不到50%的蛋白质具有足够精确的结构数据,可以理解编码蛋白质的基因变化对结构的影响。这种情况在2021年发生了巨大变化,当时DeepMind的AlphaFold人工智能系统成功地预测了与实验表征蛋白质质量相当的蛋白质结构。2022年8月,DeepMind发布了所有已知蛋白质的bb0.214亿个蛋白质结构。虽然最近的分析表明,在某些情况下,AlphaFold模型不够精确,无法进行详细的研究,很大程度上是因为进行预测所需的数据仍然过于稀少,但AlphaFold数据仍然大量增加了高质量结构数据的数量,可用于理解蛋白质的功能机制。鉴定蛋白质的组成结构域并非易事。该项目将利用强大的人工智能技术更准确地预测领域边界。初步研究已经显示出显著的改善。我们将采用由两个世界知名的蛋白质结构域分类团队(ECOD, CATH)独立开发的多个结构域检测算法,这两个团队在成功的自动化结构域检测方面都有长期的记录。他们的方法采用互补策略,可以结合起来给出共识预测,其中任务中的一致性反映了更高的置信度。另一个主要挑战将是处理数据的规模。即使考虑到由于模型质量差而造成的50%的损失,这些数据也比这些进化资源中已经分类的数据增加了100到200倍。现有的领域分配和分类管道(3D-SCAFOLD)用于整合来自两个资源(SCOP, CATH)的实验领域数据,将被重新设计为包含ECOD(比SCOP更全面),并捕获来自AlphaFold的大量预测数据。这将需要新的和更有效的工作流程来并行处理这些过程。此外,管道将更加复杂,因为需要额外的步骤来确定模型质量并删除不良模型。我们还将调整对网页和api的访问,以允许用户请求目标子集,并根据数据规模的增加执行更复杂的查询。此外,我们预计许多大的、更复杂的多结构域蛋白将非常具有挑战性,导致不同资源提供的结果之间存在差异。我们将举办研讨会,让团队就共识任务达成一致。为了应对数据的规模,我们将首先瞄准致病生物中的蛋白质、对粮食安全至关重要的作物以及与人类健康和福祉相关的蛋白质家族,包括对环境修复和生产具有商业价值的化合物至关重要的酶家族。
英文摘要
Proteins play a major role in most important processes in life, such as the digestion of nutrients, immune response, and cellular regulation. They are comprised of long polymers that fold into compact globular forms known as domains. Most proteins have at least two domains and some are composed of dozens. Domains tend to be associated with specific functions, although sometimes an important function will result from combining multiple domains. 3D structure data and models are particularly valuable for detecting the pockets and surface features linked to domain function. Determining the structure and orientations of the constituent domains is important for understanding the overall function of the protein and the dynamic conformational changes linked to that. Until recently, structural data for proteins was very sparse, with <1% of all known proteins experimentally characterised. Whilst structures can be predicted with reasonable accuracy when the structure of a close relative is known, for a significant proportion of proteins such data did not exist. Even for important organisms like humans or wheat, <50% of proteins had structural data accurate enough to understand the structural impacts of changes in the genes coding the proteins.This situation changed dramatically in 2021 when DeepMind's AlphaFold AI system succeeded in predicting protein structures of comparable quality to experimentally characterised proteins. In August 2022, DeepMind released >214 million protein structures for all known proteins. Whilst recent analyses showed that in some cases AlphaFold models are not accurate enough for detailed studies, largely because the data needed to make the prediction is still too sparse, the AlphaFold data still massively increases the amount of high-quality structural data available for understanding the mechanisms by which proteins function.Identifying constituent domains in a protein is not trivial. This project will exploit powerful AI technologies to more accurately predict domain boundaries. Preliminary studies are already showing significant improvements. We will apply multiple domain detection algorithms independently developed by two world-renowned protein domain classification teams (ECOD, CATH), both of whom have long track records in successfully automating domain detection. Their methods employ complementary strategies that can be combined to give a consensus prediction where agreement in assignments reflects higher confidence levels. Another major challenge will be coping with the scale of the data. Even allowing for a 50% loss due to poor model quality, the data represents a >200-fold increase in the data already classified in these evolutionary resources. An existing domain assignment and classification pipeline (3D-SCAFOLD) built to integrate experimental domain data from two resources (SCOP, CATH) will be re-engineered to incorporate ECOD (which is much more comprehensive than SCOP) and capture the vast predicted data from AlphaFold. This will require new and more efficient workflows that parallelise the processes. Furthermore, the pipeline will be more complex as additional steps will be necessary to determine the model quality and remove poor models. We will also adapt access to the webpages and APIs to allow users to request targeted subsets and perform more complex queries needed by the increase in the scale of the data.In addition, we expect that many large, more complex multidomain proteins will be very challenging, leading to discrepancies between the results provided by the different resources. We will hold workshops for the teams to agree on consensus assignments.To cope with the scale of the data, we will initially target proteins in pathogenic organisms, crops essential for food security, and protein families linked to human health and well-being, including enzyme families important for environmental remediation and the production of commercially valuable compounds.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
ProtFunAI: AI based methods for functional annotation of proteins in crop genomes
-
批准号:BB/Y514044/1
-
项目类别:Research Grant
-
资助金额:$32.43万
-
财政年份:2024
-
负责人:Christine Orengo
-
依托单位:
Improving accuracy, coverage, and sustainability of functional protein annotation in InterPro, Pfam and FunFam using Deep Learning methods PID 7012435
-
批准号:BB/X018563/1
-
项目类别:Research Grant
-
资助金额:$16.68万
-
财政年份:2024
-
负责人:Christine Orengo
-
依托单位:
Transforming the Structural Landscape of CATH to Aid Variant Analyses in Human and Agricultural Organisms and their Pathogens
-
批准号:BB/W018802/1
-
项目类别:Research Grant
-
资助金额:$111.5万
-
财政年份:2022
-
负责人:Christine Orengo
-
依托单位:
Unlocking the chemical potential of plants: Predicting function from DNA sequence for complex enzyme superfamilies
-
批准号:BB/V014722/1
-
项目类别:Research Grant
-
资助金额:$39.23万
-
财政年份:2022
-
负责人:Christine Orengo
-
依托单位:
CATH-FunVar - Predicting Viral and Human Variants Affecting COVID-19 Susceptibility and Severity and Repurposing Therapeutics
-
批准号:BB/W003368/1
-
项目类别:Research Grant
-
资助金额:$14.89万
-
财政年份:2021
-
负责人:Christine Orengo
-
依托单位:
3D-Gateway - Gateway to protein structure and function
-
批准号:BB/S020144/1
-
项目类别:Research Grant
-
资助金额:$37.37万
-
财政年份:2020
-
负责人:Christine Orengo
-
依托单位:
Exploiting data driven computational approaches for understanding protein structure and function in InterPro and Pfam
-
批准号:BB/S020039/1
-
项目类别:Research Grant
-
资助金额:$3.42万
-
财政年份:2020
-
负责人:Christine Orengo
-
依托单位:
SENSE - Screening of ENvironmental SEquences to discover novel protein functions, using informatics target selection and high-throughput validation
-
批准号:BB/T002735/1
-
项目类别:Research Grant
-
资助金额:$29.22万
-
财政年份:2020
-
负责人:Christine Orengo
-
依托单位:
BBSRC-NSF/BIO Expanding the fold library in the twilight zone to facilitate structure determination of macromolecular machines
-
批准号:BB/S016007/1
-
项目类别:Research Grant
-
资助金额:$43.85万
-
财政年份:2020
-
负责人:Christine Orengo
-
依托单位:
Increasing the Coverage and Accuracy of CATH for Comparative Genomics and Variant Interpretation
-
批准号:BB/R014892/1
-
项目类别:Research Grant
-
资助金额:$79.16万
-
财政年份:2018
-
负责人:Christine Orengo
-
依托单位:
FunPDBe - Community driven enrichment of PDB data with structural and functional annotations
-
批准号:BB/P023940/1
-
项目类别:Research Grant
-
资助金额:$13.34万
-
财政年份:2017
-
负责人:Christine Orengo
-
依托单位:
Expanding Genome3D and disseminating the structural annotations via InterPro and PDBe
-
批准号:BB/N019253/1
-
项目类别:Research Grant
-
资助金额:$49.25万
-
财政年份:2016
-
负责人:Christine Orengo
-
依托单位:
CATH-FunL: Improving Gene Target Selection by Predicting Functional Modules in Biological Systems
-
批准号:BB/M020088/1
-
项目类别:Research Grant
-
资助金额:$14.42万
-
财政年份:2015
-
负责人:Christine Orengo
-
依托单位:
An Greatly Expanded CATH-Gene3D with Functional Fingerprints to Characterise Proteins
-
批准号:BB/K020013/1
-
项目类别:Research Grant
-
资助金额:$78.03万
-
财政年份:2014
-
负责人:Christine Orengo
-
依托单位:
GENOME-3D: a UK network providing structure-based annotations for genotype to phenotype studies
-
批准号:BB/I025050/1
-
项目类别:Research Grant
-
资助金额:$37.5万
-
财政年份:2012
-
负责人:Christine Orengo
-
依托单位:
Exploiting High Performance Computing to Provide Functional Annotations via CATH-Gene3D
-
批准号:BB/H02364X/1
-
项目类别:Research Grant
-
资助金额:$13.88万
-
财政年份:2010
-
负责人:Christine Orengo
-
依托单位:
An Integrated CATH Resource for the Postgenomic Era
-
批准号:BB/F010451/1
-
项目类别:Research Grant
-
资助金额:$104.01万
-
财政年份:2008
-
负责人:Christine Orengo
-
依托单位:
国内基金
海外基金
登录
查看更多内容
SYNJ1蛋白片段通过促进突触蛋白NSF聚集在帕金森病发生中的机制研究
-
批准号:--
-
项目类别:青年科学基金项目
-
资助金额:30万元
-
批准年份:2022
-
负责人:邹利
-
依托单位:
NSF蛋白亚硝基化修饰所介导的GluA2 containing-AMPA受体膜稳定性在卒中后抑郁中的作用及机制研究
-
批准号:82071300
-
项目类别:面上项目
-
资助金额:55.0万元
-
批准年份:2020
-
负责人:方琪
-
依托单位:
参加中美(NSFC-NSF)生物多样性项目评审会
-
批准号:--
-
项目类别:国际(地区)合作与交流项目
-
资助金额:2万元
-
批准年份:2019
-
负责人:贺金生
-
依托单位:
参加中美(NSFC-NSF)生物多样性项目评审会
-
批准号:31981220281
-
项目类别:国际(地区)合作与交流项目
-
资助金额:2.3万元
-
批准年份:2019
-
负责人:张全发
-
依托单位:
中美(NSFC-NSF)EEID联合评审会
-
批准号:--
-
项目类别:国际(地区)合作与交流项目
-
资助金额:2.6万元
-
批准年份:2019
-
负责人:肖立华
-
依托单位:
中美(NSFC-NSF)EEID联合评审会
-
批准号:81981220037
-
项目类别:国际(地区)合作与交流项目
-
资助金额:2.1万元
-
批准年份:2019
-
负责人:段广才
-
依托单位:
中美(NSFC-NSF)EEID联合评审会
-
批准号:--
-
项目类别:国际(地区)合作与交流项目
-
资助金额:1.2万元
-
批准年份:2019
-
负责人:王四宝
-
依托单位:
Mon1b 协同NSF调控早期内吞体膜融合的机制研究
-
批准号:31671397
-
项目类别:面上项目
-
资助金额:67.0万元
-
批准年份:2016
-
负责人:李红昌
-
依托单位: