An Greatly Expanded CATH-Gene3D with Functional Fingerprints to Characterise Proteins
An Greatly Expanded CATH-Gene3D with Functional Fingerprints to Characterise Proteins
批准号:
BB/K020013/1
负责人:
Christine Orengo
金额:
$78.03万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2014
资助国家:
英国
项目状态:
已结题
起止时间:
2014 至 --
中文摘要
数以百万计的蛋白质正在被测序,它们没有已知的功能。新的CATH方法将预测它们的功能。虽然其他资源也可以预测功能,但CATH- gene3d(以下简称CATH)提供了与功能相关的结构保守特征的独特信息。结构数据揭示了蛋白质如何发挥其功能,以及当蛋白质被突变或其他遗传变异修饰时,为什么功能会发生变化。蛋白质功能信息是理解生物系统的关键,并由此延伸到药物设计、蛋白质工程和疾病。CATH是一个世界领先的资源,它将从同一祖先蛋白质进化而来的蛋白质分类为进化家族。目前,CATH将1500万个蛋白质结构域分为2600个家族。家族数据是有价值的,因为进化亲属(称为同源物)往往具有相似的三维结构和执行相似的功能。因此,CATH的好处是能够推断同源物之间的性质。这一点很重要,因为目前已知的数以百万计的蛋白质(大约2000万)中,只有不到5%的蛋白质通过实验确定了功能。即使在我们最感兴趣的生物体中,人类,也只有不到10%的蛋白质具有已知的功能。因为对蛋白质进行表征是缓慢而昂贵的,所以不可能对所有这些蛋白质进行实验研究。因此,生物学家使用CATH来预测基于其所属家族的蛋白质的功能。另一个事实是,蛋白质是由“结构域”组成的——平均每个蛋白质有两个结构域。它们是独立折叠的实体,共同作用赋予整个蛋白质的功能。CATH在结构域水平上对蛋白质进行分类,目前对自然界中发现的约70%的结构域进行分类。结构域是蛋白质的组成部分——几千个结构域以不同的方式组合在一起,形成了自然界中2000万个甚至更多的蛋白质。我们小组开发预测域函数的方法。这使得整个蛋白质的功能可以从其组成结构域的功能推断出来。因此,可以对由任何结构域组合而成的蛋白质提出功能建议。CATH使用域的三维结构信息来给出更准确的家族分类,因为结构在进化过程中比序列更高度保守。更重要的是,结构可以揭示蛋白质如何发挥其功能,以及如果在特定位点发生突变,蛋白质是否会失去其功能。我们将把CATH扩大100%。由于手动验证非常耗时,我们将开发更好的方法来自动识别远距离同源物。我们将持续发布数据(CATH-B),在人工管理之前,以便生物学家可以更快地从信息中受益。我们将与其他主要的结构分类SCOP合作,制定共同的分类策略,并提供有关家庭的补充信息。提高家族间功能遗传的准确性。我们需要这样做,因为在一些家庭中,特别是在自然界中出现频率更高的家庭中,某些亲属的功能可能会发生变化。我们将通过描述领域中的重要位置来提高准确性,这些位置在功能相似的亲属中是保守的。我们可以建立这些位置的模式,以识别共享这些模式并可能具有类似功能的其他域。我们将使生物学家更容易使用我们的网络搜索工具来确定蛋白质是否属于这些功能家族之一。我们将把它设置在云上,这样生物学家就可以用他们使用新的测序技术获得的大量数据集快速搜索CATH。这些技术捕获在不同条件下表达的蛋白质。我们的网页将报道它们的功能和蛋白质的变化,这些变化可能会改变导致疾病的功能
英文摘要
There are millions of proteins being sequenced which have no known function. New CATH methods will predict their functions. Whilst other resources do also predict function, CATH-Gene3D (referred to below as CATH) provides unique information on structurally conserved features linked to function. Structure data reveals how proteins perform their function and why the function changes if the protein is modified by mutations or other genetic variations. Protein function information is key to understanding biological systems and by extension drug design, protein engineering and disease.CATH is a world leading resource that classifies proteins evolved from the same ancestral protein, into evolutionary families. Currently, CATH classifies 15 million protein domains into 2600 families. Family data is valuable because evolutionary relatives (called homologues) tend to have similar 3D structures and perform similar functions. Thus the benefit of CATH is the ability to infer properties between homologues.This is important because of the millions of proteins currently known (>20 million) less than 5% have experimentally determined functions. Even in the organism of greatest interest to us, human, <10% of proteins have known functions. Because it can be slow and very expensive to characterise proteins it will not be possible to experimentally study all these proteins. Therefore, biologists use CATH to predict the function of a protein based on the family to which it belongs.Another fact is that proteins are made up of 'domains' - on average two per protein. These are independently folded entities that act together to confer the function of the whole protein. CATH classifies proteins at the level of the domain and currently classifies ~70% of domains found in nature. Domains are the building blocks of proteins - a few thousand of them are combined in different ways to give the 20 million proteins, or more, in nature. Our group develops methods for predicting domain functions. This allows functions of whole proteins to be deduced from the functions of their constitutive domains. Thus functions can be suggested for proteins made from any combination of domains.CATH uses information on the 3D structure of the domain to give more accurate family classifications, as structure is more highly conserved, during evolution, than the sequence. Even more important - structure can reveal how the protein performs its function and whether the protein loses its function if a mutation occurs at a particular site.We will expand CATH by 100%. Since manual validation is very time consuming, we will develop better methods for automatically recognising distant homologues. We will continuously release data (CATH-B), prior to manual curation, so that biologists can benefit from the information much sooner.We will collaborate with the other major structure classification SCOP to develop common classification strategies and provide complementary information on families.We will improve the accuracy of functional inheritance across a family. We need to do this because in some families, especially those occurring more frequently in nature, the functions can change in some relatives.We will improve accuracy by characterising important positions in the domain, conserved across functionally similar relatives. We can build patterns of these positions to recognise other domains sharing such patterns and likely to have similar functions.We will make it easy for biologists to use our web search tool to determine if a protein belongs to one of these functional families. We will set this up on the Cloud so that biologists can quickly search CATH with the massive datasets they obtain using new sequencing technologies. These technologies capture proteins expressed under different conditions. Our web pages will report their functions and variations in the protein which could modify function causing disease
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
登录
查看更多内容
DOI:
10.1016/j.sbi.2016.06.018
发表时间:
2016-10
期刊:
CURRENT OPINION IN STRUCTURAL BIOLOGY
影响因子:
6.8
作者:
[Berman, Helen M., Burley, Stephen K., Kleywegt, Gerard J., Markley, John L., Nakamura, Haruki, Velankar, Sameer]
通讯作者:
Velankar, Sameer
DOI:
10.1016/j.gde.2015.09.005
发表时间:
2015-12
期刊:
Current opinion in genetics & development
影响因子:
4
作者:
[Das S, Dawson NL, Orengo CA]
通讯作者:
Orengo CA
DOI:
10.1007/s10822-014-9770-y
发表时间:
2014-10
期刊:
JOURNAL OF COMPUTER-AIDED MOLECULAR DESIGN
影响因子:
3.5
作者:
[Berman, Helen M., Kleywegt, Gerard J., Nakamura, Haruki, Markley, John L.]
通讯作者:
Markley, John L.
DOI:
10.1093/bioinformatics/btv398
发表时间:
2015-11-01
期刊:
Bioinformatics (Oxford, England)
影响因子:
--
作者:
[Das S, Lee D, Sillitoe I, Dawson NL, Lees JG, Orengo CA]
通讯作者:
Orengo CA
DOI:
10.1093/nar/gkv488
发表时间:
2015-07-01
期刊:
Nucleic acids research
影响因子:
14.9
作者:
[Das S, Sillitoe I, Lee D, Lees JG, Dawson NL, Ward J, Orengo CA]
通讯作者:
Orengo CA
BBSRC-NSF/BIO: An AI-based domain classification platform for 200 million 3D-models of proteins to reveal protein evolution
-
批准号:BB/Y001117/1
-
项目类别:Research Grant
-
资助金额:$34.21万
-
财政年份:2024
-
负责人:Christine Orengo
-
依托单位:
ProtFunAI: AI based methods for functional annotation of proteins in crop genomes
-
批准号:BB/Y514044/1
-
项目类别:Research Grant
-
资助金额:$32.43万
-
财政年份:2024
-
负责人:Christine Orengo
-
依托单位:
Improving accuracy, coverage, and sustainability of functional protein annotation in InterPro, Pfam and FunFam using Deep Learning methods PID 7012435
-
批准号:BB/X018563/1
-
项目类别:Research Grant
-
资助金额:$16.68万
-
财政年份:2024
-
负责人:Christine Orengo
-
依托单位:
Transforming the Structural Landscape of CATH to Aid Variant Analyses in Human and Agricultural Organisms and their Pathogens
-
批准号:BB/W018802/1
-
项目类别:Research Grant
-
资助金额:$111.5万
-
财政年份:2022
-
负责人:Christine Orengo
-
依托单位:
Unlocking the chemical potential of plants: Predicting function from DNA sequence for complex enzyme superfamilies
-
批准号:BB/V014722/1
-
项目类别:Research Grant
-
资助金额:$39.23万
-
财政年份:2022
-
负责人:Christine Orengo
-
依托单位:
CATH-FunVar - Predicting Viral and Human Variants Affecting COVID-19 Susceptibility and Severity and Repurposing Therapeutics
-
批准号:BB/W003368/1
-
项目类别:Research Grant
-
资助金额:$14.89万
-
财政年份:2021
-
负责人:Christine Orengo
-
依托单位:
3D-Gateway - Gateway to protein structure and function
-
批准号:BB/S020144/1
-
项目类别:Research Grant
-
资助金额:$37.37万
-
财政年份:2020
-
负责人:Christine Orengo
-
依托单位:
Exploiting data driven computational approaches for understanding protein structure and function in InterPro and Pfam
-
批准号:BB/S020039/1
-
项目类别:Research Grant
-
资助金额:$3.42万
-
财政年份:2020
-
负责人:Christine Orengo
-
依托单位:
SENSE - Screening of ENvironmental SEquences to discover novel protein functions, using informatics target selection and high-throughput validation
-
批准号:BB/T002735/1
-
项目类别:Research Grant
-
资助金额:$29.22万
-
财政年份:2020
-
负责人:Christine Orengo
-
依托单位:
BBSRC-NSF/BIO Expanding the fold library in the twilight zone to facilitate structure determination of macromolecular machines
-
批准号:BB/S016007/1
-
项目类别:Research Grant
-
资助金额:$43.85万
-
财政年份:2020
-
负责人:Christine Orengo
-
依托单位:
Increasing the Coverage and Accuracy of CATH for Comparative Genomics and Variant Interpretation
-
批准号:BB/R014892/1
-
项目类别:Research Grant
-
资助金额:$79.16万
-
财政年份:2018
-
负责人:Christine Orengo
-
依托单位:
FunPDBe - Community driven enrichment of PDB data with structural and functional annotations
-
批准号:BB/P023940/1
-
项目类别:Research Grant
-
资助金额:$13.34万
-
财政年份:2017
-
负责人:Christine Orengo
-
依托单位:
Expanding Genome3D and disseminating the structural annotations via InterPro and PDBe
-
批准号:BB/N019253/1
-
项目类别:Research Grant
-
资助金额:$49.25万
-
财政年份:2016
-
负责人:Christine Orengo
-
依托单位:
CATH-FunL: Improving Gene Target Selection by Predicting Functional Modules in Biological Systems
-
批准号:BB/M020088/1
-
项目类别:Research Grant
-
资助金额:$14.42万
-
财政年份:2015
-
负责人:Christine Orengo
-
依托单位:
GENOME-3D: a UK network providing structure-based annotations for genotype to phenotype studies
-
批准号:BB/I025050/1
-
项目类别:Research Grant
-
资助金额:$37.5万
-
财政年份:2012
-
负责人:Christine Orengo
-
依托单位:
Exploiting High Performance Computing to Provide Functional Annotations via CATH-Gene3D
-
批准号:BB/H02364X/1
-
项目类别:Research Grant
-
资助金额:$13.88万
-
财政年份:2010
-
负责人:Christine Orengo
-
依托单位:
An Integrated CATH Resource for the Postgenomic Era
-
批准号:BB/F010451/1
-
项目类别:Research Grant
-
资助金额:$104.01万
-
财政年份:2008
-
负责人:Christine Orengo
-
依托单位:
海外基金