课题基金 / 基金详情

3D-Gateway - Gateway to protein structure and function

3D-Gateway - Gateway to protein structure and function
3D-Gateway - 蛋白质结构和功能的门户
批准号:
BB/S020144/1
负责人:
Christine Orengo
金额:
$37.37万
依托单位:
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2020
资助国家:
英国
项目状态:
已结题
起止时间:
2020 至 --

项目摘要

项目成果

Christine Orengo的其他基金

相似基金

相关文献

中文摘要
翻译
蛋白质由有机分子长链组成,折叠成紧凑的球形三维结构。了解这种结构可以对裂缝、口袋或其他对结合细胞中其他分子(如小分子或蛋白质)很重要的表面特征提供非常有价值的见解。对结构的了解对于设计结合这些特征并抑制蛋白质的药物也是必不可少的,也有助于了解蛋白质残基的突变是否会影响其稳定性或功能,从而导致疾病。通过实验确定蛋白质的结构是具有挑战性的,这就是为什么只有一小部分已知蛋白质(1.2亿蛋白质中的14.5万)被表征出来的原因。然而,已经开发出强大的计算方法,通过继承已知结构的进化相关蛋白质的结构信息来预测蛋白质结构。最近,这些预测技术变得更加强大,因为人们发现了利用进化数据的新方法,可以更准确地限制蛋白质中的接触。应用这些技术,可以预测大部分未表征蛋白质的结构。例如,对于人类蛋白质,大约5%的结构是已知的,但另外88%的结构可以建模,其中一些非常精确,从而为设计治疗人类疾病的药物提供了重要的框架。当继承远亲之间的结构数据时,人们必须更加谨慎,大多数预测方法都会为所产生的模型返回置信度分数。该项目将建立一个基础设施(3D-Beacons),将实验确定的结构与应用不同算法的团队生成的预测结构聚合在一起。这将对来自与粮食安全和人类健康有关的选定生物体的蛋白质进行检测——其中一些将是威胁人类或动物/作物的致病菌。我们将使用这些数据来注释UniProt资源中的蛋白质,该资源每月被超过75万的独立用户广泛使用。由于预测方法存在于许多不同的实验室,通过这种方式汇集数据,我们可以显着增加具有结构数据的蛋白质数量。此外,结合由独立算法构建的模型,我们可以比较3d模型,以找出哪些部分是一致的,而不管使用哪种方法,哪些部分在不同的方法之间存在差异,显然难以建模。因此,我们将使用这些汇总数据来研究计算蛋白质中每个位置的模型质量的最佳策略。我们将建立网页,以显示已知和预测的结构,为一个给定的蛋白质。确定整个蛋白质的结构可能很困难,因此,在适当的情况下,我们将同时显示实验结构和预测结构,并非常小心地标记结构的来源信息(例如使用的方法)和数据的可靠性(例如置信度)。我们还将使用我们的3D-Beacons基础设施来汇总蛋白质结构上已知和预测的功能位点的信息,并在网页上显示这些数据,以及来源和置信度的信息。将位点数据映射到结构上将特别有助于制定规则,使我们能够衡量没有实验表征的蛋白质是否与具有实验表征的进化相关蛋白质具有相同的功能。具有相同功能的亲缘关系应该具有相同的关键功能位点残基。有了这些规则,我们将能够为UniProt中的数百万个蛋白质提供结构和功能注释。新的数据将代表具有结构和功能位点信息的UniProt序列数量增加十倍或更多。UniProt也被工业研究人员广泛使用,因此信息的扩展将产生非常重大的影响。
英文摘要
Proteins comprise long chains of organic molecules that fold into compact globular 3-dimensional structures. Knowing this structure can give very valuable insights into the clefts, pockets or other surface features important for binding other molecules in the cell eg small molecules or proteins. Knowledge of the structure is also essential for designing drugs that bind to these features and inhibit the protein and can also help in understanding whether mutations in the protein's residues affect its stability or function, leading to disease. Experimentally determining the structure can be challenging, which is why only a small percentage of known proteins (~145,000 out of 120 million) have been characterised. However, powerful computational methods have been developed that predict protein structures by inheriting structural information from evolutionary related proteins whose structures are known. These prediction techniques have been made even more powerful, recently, as new ways of exploiting the evolutionary data have been found that more accurately constrain contacts in the protein. Applying these techniques, structures can be predicted for a large proportion of uncharacterised proteins. For example, for human proteins about 5% of the structures are known but a further 88% can be modelled, some to very high accuracy, thereby providing important frameworks for designing drugs to treat human diseases. When inheriting structural data between distant relatives one has to be much more cautious and most prediction methods return a confidence score for the models produced. This project will build an infrastructure (3D-Beacons) that aggregates experimentally determined structures with predicted structures generated by groups applying different algorithms. This will be done for proteins from selected organisms relevant to food security and human health - some will be pathogenic bacteria that threaten humans or animals/crops. We will use this data to annotate proteins in the UniProt resource, widely used by more than 750,000 unique users each month. Since the prediction methods reside in many different labs, by pooling the data in this way we can significantly increase the number of proteins with structural data. In addition, combining models built by independent algorithms allows us to compare 3D-models to find which parts agree regardless of method and which parts vary between methods and are clearly harder to model. Therefore, we will use this aggregated data to research the best strategies for calculating model quality at each position in the protein.We will build web pages to display the known and predicted structures for a given protein. It can be difficult to determine the structure of the whole protein so, where appropriate, we will display both experimental and predicted structures, taking great care to label the structures with information on the source (eg method used) and reliability of the data (eg confidence).We will also use our 3D-Beacons infrastructure to aggregate information on known and predicted functional sites on the protein structure and display this data on web pages, together with information on source and confidence. The site data mapped onto structure will be particularly helpful for developing rules that allow us to gauge whether a protein with no experimental characterisation has the same function as an evolutionary related protein with experimental characterisation. Relatives sharing the same function should have the same key functional site residues. With these rules we will be able to provide structural and functional annotations for millions of proteins in UniProt. The new data will represent a tenfold or more increase in the number of UniProt sequences which have structural and functional site information. UniProt is also widely used by researchers in industry and thus this expansion in information will have a very significant impact.
期刊论文(8)
专著(0)
科研奖励(0)
会议论文
DOI: 10.1093/bib/bbaa362
发表时间: 2021-03-22
期刊: Briefings in bioinformatics
影响因子: 9.5
作者: [Waman VP, Sen N, Varadi M, Daina A, Wodak SJ, Zoete V, Velankar S, Orengo C]
通讯作者: Orengo C
DOI: 10.1016/j.jmb.2023.168021
发表时间: 2023-06-24
期刊: JOURNAL OF MOLECULAR BIOLOGY
影响因子: 5.6
作者: [Vallat, Brinda, Tauriello, Gerardo, Westbrook, John D.]
通讯作者: Westbrook, John D.
DOI: 10.1038/s42003-023-04488-9
发表时间: 2023-02-08
期刊: Communications biology
影响因子: 5.9
作者: []
通讯作者:
CATHe: Detection of remote homologues for CATH superfamilies using embeddings from protein language models
CATHe:使用蛋白质语言模型的嵌入检测 CATH 超家族的远程同源物
DOI: 10.1101/2022.03.10.483805
发表时间: 2022
期刊:
影响因子: --
作者: [Nallapareddy V]
通讯作者: Nallapareddy V
共 6 条
    BBSRC-NSF/BIO: An AI-based domain classification platform for 200 million 3D-models of proteins to reveal protein evolution
    • 批准号:
      BB/Y001117/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $34.21万
    • 财政年份:
      2024
    • 负责人:
      Christine Orengo
    • 依托单位:
    ProtFunAI: AI based methods for functional annotation of proteins in crop genomes
    • 批准号:
      BB/Y514044/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $32.43万
    • 财政年份:
      2024
    • 负责人:
      Christine Orengo
    • 依托单位:
    Improving accuracy, coverage, and sustainability of functional protein annotation in InterPro, Pfam and FunFam using Deep Learning methods PID 7012435
    • 批准号:
      BB/X018563/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $16.68万
    • 财政年份:
      2024
    • 负责人:
      Christine Orengo
    • 依托单位:
    Transforming the Structural Landscape of CATH to Aid Variant Analyses in Human and Agricultural Organisms and their Pathogens
    • 批准号:
      BB/W018802/1
    • 项目类别:
      Research Grant
    • 资助金额:
      $111.5万
    • 财政年份:
      2022
    • 负责人:
      Christine Orengo
    • 依托单位:
    国内基金
    海外基金
    应用Gateway技术构建Survivin、TGFβ3和TIMP-1三克隆系统并转染抑制椎间盘退变
    • 批准号:
      81171758
    • 项目类别:
      面上项目
    • 资助金额:
      58.0万元
    • 批准年份:
      2011
    • 负责人:
      陈伯华
    • 依托单位: