Expanding Genome3D and disseminating the structural annotations via InterPro and PDBe
Expanding Genome3D and disseminating the structural annotations via InterPro and PDBe
批准号:
BB/N019253/1
负责人:
Christine Orengo
金额:
$49.25万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2016
资助国家:
英国
项目状态:
已结题
起止时间:
2016 至 --
中文摘要
蛋白质的结构决定了它与其他蛋白质相互作用的方式,以及它是否或如何结合和改变它所接触的化合物。了解蛋白质的结构有助于合理化它发挥生物作用的机制。这对于理解遗传变化(如组成蛋白质的残基突变)如何破坏或改变其发挥作用的方式也很重要。被称为下一代测序的生物学革命性新技术现在允许生物学家收集大量的遗传变异数据。例如,从患有癌症或心脏病等不同疾病的人身上收集的蛋白质序列变化的信息。或者,来自在农业环境中重要的物种的蛋白质序列。例如,不同的小麦品种可能更抗霜冻或产量更高,但是,确定蛋白质的3D结构比确定其序列要困难得多,也更昂贵。对于人类、小鼠、鸡、植物和其他我们需要研究以了解疾病或确保粮食安全的真核生物来说,这尤其困难。目前,平均不到15%的蛋白质从这些重要的模式生物具有实验确定的3D结构。为了解决结构数据的这种缺陷,已经开发了用于预测蛋白质结构的算法。最成功的方法是通过利用进化相关蛋白质之间已知的结构特征保守性来识别具有已知结构的亲属并继承3D信息。生成此类注释的五个世界领先的顶级资源位于英国(SUPERFAMILY,Gene 3D,Phyre,Fugure,pDomTHREADER)。这些利用SCOP和CATH结构分类中的结构相关物-这两个世界领先的资源捕获域结构的信息-用作预测未表征的相关物结构的模板。Genome 3D资源于2012年推出,整合了来自所有五个资源的结构域预测,用于研究生物系统的十种模式生物,对人类健康(例如人类,小鼠)或农业和粮食安全(例如植物)的研究非常重要。虽然资源使用的算法对于识别非常遥远的关系和继承亲属之间的结构信息非常强大,但它们的准确率<90%。然而,通过将所有数据组合在一个资源中,并确定所有方法一致的蛋白质位置,可以提供更可靠的注释。如果在SCOP和CATH中已经确定了等价的亲属(即家族),那么找到这些共有区域就更容易了,因此该项目的很大一部分涉及这些资源之间的映射。我们现在希望继续这个项目,改进SCOP和CATH的映射,并使用它来增加Genome 3D提供的可靠的共有数据量。我们将包括对健康和农业重要的其他生物。然而,该项目的一个主要好处是将Genome 3D结构数据与InterPro中结构未表征的序列整合在一起,InterPro是一个世界领先的资源,它结合了来自全球11个不同资源的蛋白质家族信息。通过在InterPro中包含家族的Genome 3D数据,我们将能够将我们可以提供结构数据的蛋白质数量增加十倍。此外,我们将提供一个非常直观的基于Web的查看器,用于查看结构并评估序列中任何变化对蛋白质功能的可能影响。由于许多生物学家不熟悉结构数据在评估遗传变异方面的价值,我们将开发基于网络的培训材料,并在我们的研究所和国际会议上安排研讨会。
英文摘要
The structure of a protein dictates the manner in which it interacts with other proteins and whether or how it binds and changes the compounds it is exposed to. Knowing a protein's structure can help rationalise the mechanism by which it performs its biological role. It is also important for understanding how genetic changes such as mutations in the residues that make up the protein, can destroy or modify the way in which it performs that role. Revolutionary new technologies in biology, known as next generation sequencing, are now allowing biologists to collect vast amounts of genetic variation data. For example, information on changes in the sequences of proteins collected from humans suffering from different diseases like cancer or heart disease. Alternatively, sequences of proteins from species important in an agricultural context. For example different strains of wheat that may be more resistant to frost or produce higher yields.However, it is much harder and more expensive to determine the 3D structure of a protein than its sequence. It is particularly difficult for human, mouse, chicken, plants and other eukaryotic organisms that we need to study to understand disease or ensure food security. Currently, on average less than 15% of proteins from these important model organisms have an experimentally determined 3D structure. To address this deficit of structural data, algorithms have been developed for predicting the structure of a protein. The most successful approaches identify a relative having a known structure and inherit 3D information by exploiting the known conservation of structural features between evolutionary related proteins. Five of the top world-leading resources generating such annotations are based in the UK (SUPERFAMILY, Gene3D, Phyre, Fugure, pDomTHREADER). These exploit structural relatives in the SCOP and CATH structural classification - the two world leading resources capturing information on domain structures - to use as templates for predicting structures of uncharacterised relatives. The Genome3D resource, which was launched in 2012, integrates domain structure predictions from all five resources for ten model organisms used to study biological systems and important for the study of human health (e.g. human, mouse) or agriculture and food security (e.g. plant). Although the algorithms used by the resources are powerful for recognising even very remote relationships and inheriting structural information between relatives, their accuracy is < 90%. However, by combining all the data in a single resource and identifying positions in the protein where all the methods agree, it is possible to provide much more reliable annotations. Since it is easier to find these consensus regions if equivalent sets of relatives (i.e. families) in SCOP and CATH have been identified, a large part of the project involves mapping between these resources.We now wish to continue this project, improving the mapping of SCOP and CATH and using this to increase the amount of reliable consensus data that Genome3D provides. We will include additional organisms important for health and agriculture. However, a major benefit from this project will be the integration of the Genome3D structural data with structurally uncharacterised sequences in InterPro, a world-leading resource that combines information on protein families from 11 different resources worldwide. By including Genome3D data for families in InterPro we will be able to increase the number of proteins for which we can provide structural data ten-fold. In addition we will provide a very intuitive web-based viewer for looking at the structures and assessing the likely impacts of any changes in the sequence on the function of the protein. Since many biologists are unfamiliar with the value of structural data in assessing genetic variations we will develop web-based training material and arrange workshops both in our institutes and at international meetings.
期刊论文(6)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1093/nar/gky1100
发表时间:
2019-01-08
期刊:
Nucleic acids research
影响因子:
14.9
作者:
[Mitchell AL, Attwood TK, Babbitt PC, Blum M, Bork P, Bridge A, Brown SD, Chang HY, El-Gebali S, Fraser MI, Gough J, Haft DR, Huang H, Letunic I, Lopez R, Luciani A, Madeira F, Marchler-Bauer A, Mi H, Natale DA, Necci M, Nuka G, Orengo C, Pandurangan AP, Paysan-Lafosse T, Pesseat S, Potter SC, Qureshi MA, Rawlings ND, Redaschi N, Richardson LJ, Rivoire C, Salazar GA, Sangrador-Vegas A, Sigrist CJA, Sillitoe I, Sutton GG, Thanki N, Thomas PD, Tosatto SCE, Yong SY, Finn RD]
通讯作者:
Finn RD
DOI:
10.1093/nar/gky1097
发表时间:
2019-01-08
期刊:
Nucleic acids research
影响因子:
14.9
作者:
[Sillitoe I, Dawson N, Lewis TE, Das S, Lees JG, Ashford P, Tolulope A, Scholes HM, Senatorov I, Bujan A, Ceballos Rodriguez-Conde F, Dowling B, Thornton J, Orengo CA]
通讯作者:
Orengo CA
DOI:
10.1093/nar/gkw1098
发表时间:
2017-01-04
期刊:
Nucleic acids research
影响因子:
14.9
作者:
[Dawson NL, Lewis TE, Das S, Lees JG, Lee D, Ashford P, Orengo CA, Sillitoe I]
通讯作者:
Sillitoe I
DOI:
10.1093/nar/gkw1107
发表时间:
2017-01-04
期刊:
Nucleic acids research
影响因子:
14.9
作者:
[Finn RD, Attwood TK, Babbitt PC, Bateman A, Bork P, Bridge AJ, Chang HY, Dosztányi Z, El-Gebali S, Fraser M, Gough J, Haft D, Holliday GL, Huang H, Huang X, Letunic I, Lopez R, Lu S, Marchler-Bauer A, Mi H, Mistry J, Natale DA, Necci M, Nuka G, Orengo CA, Park Y, Pesseat S, Piovesan D, Potter SC, Rawlings ND, Redaschi N, Richardson L, Rivoire C, Sangrador-Vegas A, Sigrist C, Sillitoe I, Smithers B, Squizzato S, Sutton G, Thanki N, Thomas PD, Tosatto SC, Wu CH, Xenarios I, Yeh LS, Young SY, Mitchell AL]
通讯作者:
Mitchell AL
BBSRC-NSF/BIO: An AI-based domain classification platform for 200 million 3D-models of proteins to reveal protein evolution
-
批准号:BB/Y001117/1
-
项目类别:Research Grant
-
资助金额:$34.21万
-
财政年份:2024
-
负责人:Christine Orengo
-
依托单位:
ProtFunAI: AI based methods for functional annotation of proteins in crop genomes
-
批准号:BB/Y514044/1
-
项目类别:Research Grant
-
资助金额:$32.43万
-
财政年份:2024
-
负责人:Christine Orengo
-
依托单位:
Improving accuracy, coverage, and sustainability of functional protein annotation in InterPro, Pfam and FunFam using Deep Learning methods PID 7012435
-
批准号:BB/X018563/1
-
项目类别:Research Grant
-
资助金额:$16.68万
-
财政年份:2024
-
负责人:Christine Orengo
-
依托单位:
Transforming the Structural Landscape of CATH to Aid Variant Analyses in Human and Agricultural Organisms and their Pathogens
-
批准号:BB/W018802/1
-
项目类别:Research Grant
-
资助金额:$111.5万
-
财政年份:2022
-
负责人:Christine Orengo
-
依托单位:
Unlocking the chemical potential of plants: Predicting function from DNA sequence for complex enzyme superfamilies
-
批准号:BB/V014722/1
-
项目类别:Research Grant
-
资助金额:$39.23万
-
财政年份:2022
-
负责人:Christine Orengo
-
依托单位:
CATH-FunVar - Predicting Viral and Human Variants Affecting COVID-19 Susceptibility and Severity and Repurposing Therapeutics
-
批准号:BB/W003368/1
-
项目类别:Research Grant
-
资助金额:$14.89万
-
财政年份:2021
-
负责人:Christine Orengo
-
依托单位:
3D-Gateway - Gateway to protein structure and function
-
批准号:BB/S020144/1
-
项目类别:Research Grant
-
资助金额:$37.37万
-
财政年份:2020
-
负责人:Christine Orengo
-
依托单位:
Exploiting data driven computational approaches for understanding protein structure and function in InterPro and Pfam
-
批准号:BB/S020039/1
-
项目类别:Research Grant
-
资助金额:$3.42万
-
财政年份:2020
-
负责人:Christine Orengo
-
依托单位:
SENSE - Screening of ENvironmental SEquences to discover novel protein functions, using informatics target selection and high-throughput validation
-
批准号:BB/T002735/1
-
项目类别:Research Grant
-
资助金额:$29.22万
-
财政年份:2020
-
负责人:Christine Orengo
-
依托单位:
BBSRC-NSF/BIO Expanding the fold library in the twilight zone to facilitate structure determination of macromolecular machines
-
批准号:BB/S016007/1
-
项目类别:Research Grant
-
资助金额:$43.85万
-
财政年份:2020
-
负责人:Christine Orengo
-
依托单位:
Increasing the Coverage and Accuracy of CATH for Comparative Genomics and Variant Interpretation
-
批准号:BB/R014892/1
-
项目类别:Research Grant
-
资助金额:$79.16万
-
财政年份:2018
-
负责人:Christine Orengo
-
依托单位:
FunPDBe - Community driven enrichment of PDB data with structural and functional annotations
-
批准号:BB/P023940/1
-
项目类别:Research Grant
-
资助金额:$13.34万
-
财政年份:2017
-
负责人:Christine Orengo
-
依托单位:
CATH-FunL: Improving Gene Target Selection by Predicting Functional Modules in Biological Systems
-
批准号:BB/M020088/1
-
项目类别:Research Grant
-
资助金额:$14.42万
-
财政年份:2015
-
负责人:Christine Orengo
-
依托单位:
An Greatly Expanded CATH-Gene3D with Functional Fingerprints to Characterise Proteins
-
批准号:BB/K020013/1
-
项目类别:Research Grant
-
资助金额:$78.03万
-
财政年份:2014
-
负责人:Christine Orengo
-
依托单位:
GENOME-3D: a UK network providing structure-based annotations for genotype to phenotype studies
-
批准号:BB/I025050/1
-
项目类别:Research Grant
-
资助金额:$37.5万
-
财政年份:2012
-
负责人:Christine Orengo
-
依托单位:
Exploiting High Performance Computing to Provide Functional Annotations via CATH-Gene3D
-
批准号:BB/H02364X/1
-
项目类别:Research Grant
-
资助金额:$13.88万
-
财政年份:2010
-
负责人:Christine Orengo
-
依托单位:
An Integrated CATH Resource for the Postgenomic Era
-
批准号:BB/F010451/1
-
项目类别:Research Grant
-
资助金额:$104.01万
-
财政年份:2008
-
负责人:Christine Orengo
-
依托单位:
海外基金