课题基金 / 基金详情

18-BBSRC-NSF/BIO : CIBR:Implementing an explicit phylogenetic framework for large-scale protein sequence annotation

18-BBSRC-NSF/BIO : CIBR:Implementing an explicit phylogenetic framework for large-scale protein sequence annotation
18-BBSRC-NSF/BIO:CIBR:为大规模蛋白质序列注释实施明确的系统发育框架
批准号:
BB/T010541/1
负责人:
Maria J. Martin
金额:
$51.25万
依托单位国家:
英国
项目类别:
Research Grant
财政年份:
2020
资助国家:
英国
项目状态:
已结题
起止时间:
2020 至 --
关键词:

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
点击翻译按钮获取中文摘要
英文摘要
Proteins are the primary molecular machines that perform the instructions encoded in our genomes. Proteins ultimately shape the response of our cells, tissues, organs, and bodies to the surrounding environment, either directly (e.g. muscle contraction) or through their functional outputs (e.g. the electrical signals along the dendrites to produce a nerve impulse or action potential). Therefore, understanding the functional role(s) performed by each protein is critical to research and development in many areas of science, particularly biology, medicine and applied biotechnology. The rapid increase in throughput of next-generation sequencing technologies has important ramifications, in that our ability to sequence an organism's genome and determine the proteins it encodes far out paces our ability to experimentally characterise the function of a protein. Thus, for every functionally characterised protein, there are now many thousands of proteins that will never be experimentally characterised. Molecular biology increasingly relies on our ability to computationally group related sequences and to transfer functional annotations from the few experimentally characterised proteins, to those related, yet uncharacterised, proteins. Knowledge on proteins has been collected and stored in public databases like UniProt, a world-leading resource on protein sequences and function. Currently, there are over 150 million sequences in UniProt, with the number doubling every two years. Therefore, it is crucial to develop new and reliable computational methods for inferring protein function that can be scaled to billions of sequences. We aim to implement an annotation system that incorporates evolutionary information, permitting the level of annotation transfer to be tuned accordingly, while also ensuring scalability and speed of annotation that meets current and future demands. This new annotation system will integrate the most innovative features present in two pre-existing methods that are currently used in producing world-class resources. The Gene Ontology (GO) Consortium has developed software for explicit evolutionary modelling of GO annotation gain and loss along specific branches of phylogenetic trees, and has applied it to inferring GO annotations for experimentally uncharacterised proteins. UniProt has developed the UniRule system that applies annotation "rules" that combines information on protein families and domains (from the InterPro resource), with a range of other types of information like taxonomy, to make more precise and informative annotations. Our goal is to create a next-generation, large-scale annotation system that merges the two approaches, and to implement this annotation system in the UniProt resource, thereby increasing the quality of functional annotations in the database for the benefit of the scientific community. We propose three specific aims to achieve this goal: (1) convert existing UniRule rules into explicit evolutionary models, (2) integrate software to apply the evolutionary models (TreeGrafter) into the UniProt annotation pipeline, and (3) develop software for ongoing curation of new evolutionary models of additional annotation types and protein families. The result will be an annotation pipeline based on explicit evolutionary principles, which will enable seamless sharing of information between the UniProt and GO curation processes, and substantially improve the accuracy, comprehensiveness and informativeness of inferred protein annotations in public databases.
期刊论文(10)
专著(0)
科研奖励(0)
会议论文
DOI: 10.3390/metabo11010048
发表时间: 2021-01-12
期刊: Metabolites
影响因子: 4.1
作者: [Feuermann M, Boutet E, Morgat A, Axelsen KB, Bansal P, Bolleman J, de Castro E, Coudert E, Gasteiger E, Géhant S, Lieberherr D, Lombardot T, Neto TB, Pedruzzi I, Poux S, Pozzato M, Redaschi N, Bridge A, On Behalf Of The UniProt Consortium]
通讯作者: On Behalf Of The UniProt Consortium
DOI: 10.1093/bioinformatics/btaa485
发表时间: 2020-11-01
期刊: Bioinformatics (Oxford, England)
影响因子: --
作者: [MacDougall A, Volynkin V, Saidi R, Poggioli D, Zellner H, Hatton-Ellis E, Joshi V, O'Donovan C, Orchard S, Auchincloss AH, Baratin D, Bolleman J, Coudert E, de Castro E, Hulo C, Masson P, Pedruzzi I, Rivoire C, Arighi C, Wang Q, Chen C, Huang H, Garavelli J, Vinayaka CR, Yeh LS, Natale DA, Laiho K, Martin MJ, Renaux A, Pichler K, UniProt Consortium]
通讯作者: UniProt Consortium
DOI: 10.1093/nar/gkaa977
发表时间: 2021-01-08
期刊: Nucleic acids research
影响因子: 14.9
作者: [Blum M, Chang HY, Chuguransky S, Grego T, Kandasaamy S, Mitchell A, Nuka G, Paysan-Lafosse T, Qureshi M, Raj S, Richardson L, Salazar GA, Williams L, Bork P, Bridge A, Gough J, Haft DH, Letunic I, Marchler-Bauer A, Mi H, Natale DA, Necci M, Orengo CA, Pandurangan AP, Rivoire C, Sigrist CJA, Sillitoe I, Thanki N, Thomas PD, Tosatto SCE, Wu CH, Bateman A, Finn RD]
通讯作者: Finn RD
DOI: 10.1016/j.mcpro.2023.100591
发表时间: 2023-08
期刊: MOLECULAR & CELLULAR PROTEOMICS
影响因子: 7
作者: [Bowler-Barnett, E. H., Fan, J., Luo, J., Magrane, M., Martin, M. J., Orchard, S.]
通讯作者: Orchard, S.
7
    海外基金