Learning from biological context for protein fitness estimation and design
Learning from biological context for protein fitness estimation and design
批准号:
2593955
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2021
资助国家:
英国
项目状态:
未结题
起止时间:
2021 至 --
中文摘要
蛋白质,即氨基酸序列,是生命的基本组成部分,是细胞过程的主力。它们的不同功能,从催化化学反应到提供结构支持,支撑着生物有机体的复杂机械。进化已经对所有可能的氨基酸序列进行了大规模的实验:那些编码功能蛋白质的序列存活下来,那些不编码的氨基酸序列灭绝了。通过观察生命中一组相关的蛋白质,我们可以开始了解它的进化史,这是解锁医学、生物技术和生物学各个领域众多进步的关键。该项目属于EPSRC人工智能技术研究领域,由哈佛大学的黛比·马克斯教授共同担任顾问。计算生物学家使用越来越复杂的统计模型来分析蛋白质进化。最近,旨在揭示生命语言的大型蛋白质语言模型使这些模型扩展到整个蛋白质组成为可能。这些模型已被证明可以概括蛋白质进化或系统发育,即使相关同源物的集合很小。该项目旨在利用从背景学习和非参数建模领域到蛋白质统计建模的新方法。在上下文中,学习允许模型从上下文中学习,例如通过添加一些示例。非参数建模允许模型从显式数据点学习,而不必在其参数化权重中记忆整个数据集。通过组合这些方法,目的是更好地利用和检索蛋白质进化提供的上下文,在相同的计算成本下提高性能。这些方法将允许更好地研究不可剥夺的蛋白质序列,如无序区或抗体,以及对序列的插入和缺失进行建模。这种方法的发展既可以量化给定蛋白质序列的致病性,诊断疾病,也可以优化序列的功能,这与用于化学过程和药物开发的蛋白质生物工程有关。
英文摘要
Proteins, sequences of amino acids, are the fundamental building blocks of life, serving as the workhorses of cellular processes. Their diverse functions, from catalysing chemical reactions to providing structural support, underpin the complex machinery of living organisms. Evolution has conducted a massive experiment over the space of all possible amino acid sequences: those that encode a functional protein survive; those that don't are extinct. By looking at a set of related proteins throughout life, we can begin to understand its evolutionary history, key to unlocking numerous advancements in medicine, biotechnology, and various fields of biology. This project falls within the EPSRC Artificial Intelligence Technologies research area and is co advised by Professor Debbie Marks at Harvard University.Computational biologists have used more and more complex statistical models to analyse protein evolution. Extending the models to the whole proteome has recently been made possible by large protein language models, that aim to uncover the language of life. These models have been shown to recapitulate protein evolution, or phylogeny, even when the set of related homologs is small. The project aims to leverage new methodologies from the fields of in-context learning and non-parametric modelling to protein statistical modelling. In context learning allows models to learn from context, for example by the addition of a few examples. Nonparametric modelling allows the model to learn from explicit data points instead of having to memorize an entire dataset in its parametrised weights. By combining these methods, the aim is to better leverage and retrieve the context provided by protein evolution, to improve performance at same compute cost. These methods will allow to better study unalienable protein sequence such as disordered regions or antibodies, as well as to model insertion and deletion of sequences. The development of such method allows to both quantify pathogenicity of a given protein sequence, to diagnose disease as well as optimize a sequence for its function, with relevance to bioengineering of proteins for chemical processes and drug development.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
登录
查看更多内容
生物钟核受体Rev-erbα在缺血性卒中神经元能量代谢中的改善作用及机制研究
-
批准号:82371332
-
项目类别:面上项目
-
资助金额:49.00万元
-
批准年份:2023
-
负责人:胡琴
-
依托单位:
过表达CX45联合HCN4基因转染对起搏细胞自律性的影响
-
批准号:81170174
-
项目类别:面上项目
-
资助金额:50.0万元
-
批准年份:2011
-
负责人:周亚峰
-
依托单位:
美洲大蠊药材养殖及加工过程中化学成分动态变化与生物活性的相关性研究
-
批准号:81060329
-
项目类别:地区科学基金项目
-
资助金额:26.0万元
-
批准年份:2010
-
负责人:肖培云
-
依托单位:
慢病毒转染嵌合体HCN1+4拼接基因构建生物起搏细胞
-
批准号:81070139
-
项目类别:面上项目
-
资助金额:33.0万元
-
批准年份:2010
-
负责人:杨向军
-
依托单位:
岭南瑶区几种瑶族抗肝炎植物药的化学成分及生物活性研究
-
批准号:20772047
-
项目类别:面上项目
-
资助金额:28.0万元
-
批准年份:2007
-
负责人:岑颖洲
-
依托单位:
TB方法在有机和生物大分子体系计算研究中的应用
-
批准号:20773047
-
项目类别:面上项目
-
资助金额:26.0万元
-
批准年份:2007
-
负责人:吕文彩
-
依托单位:
天然生物材料的多尺度力学与仿生研究
-
批准号:10732050
-
项目类别:重点项目
-
资助金额:200.0万元
-
批准年份:2007
-
负责人:冯西桥
-
依托单位: