课题基金 / 基金详情

Transparent Deep Learning for Directed Protein Evolution

Transparent Deep Learning for Directed Protein Evolution
用于定向蛋白质进化的透明深度学习
批准号:
2745409
负责人:
金额:
$0.0万
依托单位:
依托单位国家:
英国
项目类别:
Studentship
财政年份:
2022
资助国家:
英国
项目状态:
未结题
起止时间:
2022 至 --

项目摘要

项目成果

相似基金

相关文献

中文摘要
翻译
蛋白质工程是一个复杂的过程,需要找到与所需功能相关的氨基酸序列。随着设计空间随着残基的数量呈指数级增长,从头设计目前是一个棘手的问题。为了克服蛋白质设计复杂性的魔咒,科学家们通常依靠一种由随机突变和选择蛋白质变体组成的迭代过程,称为定向进化(DE,1);尽管这一过程产生了显著的结果,但它极其缓慢、低通量和昂贵,因为在每一步产生功能蛋白质的概率很低。因此,在过去的30年里,科学家们开发了生物物理模型和优化方法来预测蛋白质的结构和功能;然而,这些方法通常不能扩展到大型蛋白质,并且受到潜在生物物理模型的准确性的限制。最近,机器学习(ML),特别是深度学习(DL)通过直接从数据学习与蛋白质折叠和功能相关的功能关系,在很大程度上克服了这些问题[2]。然而,理解DL模型如何进行结构和功能预测仍然是不透明和具有挑战性的[3],从而限制了它们在理解与功能蛋白质相关的生物设计原则方面的效用。目的和目标:与ZenithAI(OT/ZAI)合作,我们建议为蛋白质设计设计和构建透明和可解释的深度学习模型。蛋白质的设计空间随着所考虑的氨基酸位置的数量呈指数增长,但功能蛋白质极其罕见。因此,透明模型可以提供一种原则性的蛋白质选择方法,只需查看重要和不确定的氨基酸位置,最终减轻了蛋白质变体实验筛选的负担。工作计划。该项目由3个工作包构成。-WP1-学生将开发蛋白质工程的深度学习框架,使用最先进的变分和对抗性模型以及序列到序列模型,将使用按物种和功能分层的精选蛋白质序列信息进行培训。-WP2-学生随后将开发概率模型,通过利用模型学到的梯度和权重信息来量化设计中的不确定性,最终定义一个分数,以确定实验测试中蛋白质的优先顺序。-WP3-学生将使用该模型设计人类S1PL酶的变体,然后在实验室进行测试。S1PL是鞘磷脂途径的中心酶,对细胞的正常功能是必不可少的,它在许多疾病中都有因果作用,包括癌症和神经退行性疾病。学生将接受机器学习、统计学习和深度学习方面的培训,并将在生物序列建模和设计方面建立具有竞争力的形象。学生还将被介绍到新兴的合成生物学领域,并将学习现代DNA克隆和组装技术以及大规模蛋白质表达系统的使用。我们还非常重视可重复研究;学生将接受高级研究、软件工程和数据分析的可重复工作流程方面的培训。
英文摘要
Protein engineering is a complex process, which requires finding an amino acid sequence associated with a desired function. As the design space grows exponentially as a function of the number of residues, de-novo design is currently an intractable problem. To overcome the curse of protein design complexity, scientists routinely rely on an iterative process consisting of random mutagenesis and selection of protein variants, called Directed Evolution (DE, 1); while this process led to remarkable results, it is extremely slow, low-throughput and expensive, as the probability of generating functional proteins at each step is low. Thus, for the last 30 years, scientists have developed biophysical models and optimisation methods to predict protein structure and function in-silico; however, these methods are usually not scalable to large proteins and are limited by the accuracy of the underlying biophysical models. Recently, Machine Learning (ML) and, in particular, Deep Learning (DL) have largely overcome these problems by learning functional relationships associated with protein folding and function directly from data [2]. However, it remains opaque and challenging to understand how a DL model makes structural and functional predictions [3], thus limiting their utility in understanding the biological design principles associated with functional proteins. AIMS AND OBJECTIVES: In collaboration with ZenithAI (OT/ZAI), we propose to design and build transparent and explainable deep learning models for protein design. The protein design space increases exponentially with the number of amino acid positions considered but functional proteins are extremely rare. Therefore, transparent models can provide a principled protein selection method, by only looking at important and uncertain amino acid positions, ultimately reducing the burden of experimental screening of protein variants. WORKPLAN. The project is structured in 3 work packages. - WP1 - The student will develop a deep learning framework for protein engineering, using state-of-the-art variational and adversarial models coupled with sequence-to-sequence models, which will be trained using curated protein sequence information stratified by species and function. - WP2 - The student will then develop probabilistic models to quantify uncertainty in designs by exploiting gradient and weights information learned by the model, ultimately to define a score to prioritise proteins for experimental testing. - WP3 - The student will use the model to design variants of the human S1PL enzyme, which will then be tested in the lab. S1PL is a central enzyme in the sphingolipid pathway, which is essential for proper cell functioning and it has a causal role in many diseases, including cancer and neurodegenerative disorders.TRAINING PROGRAM. The student will receive training in machine learning, statistical learning and deep learning, and will build a competitive profile in biological sequence modelling and design. The student will be also introduced to the emerging field of synthetic biology and will learn modern DNA cloning and assembly techniques and the use of protein expression systems at scale. We also put a strong emphasis on reproducible research; the student will receive training in advanced research software engineering and in reproducible workflows for data analyses.
期刊论文(0)
专著(0)
科研奖励(0)
会议论文
国内基金
海外基金
Deep Seek引导下预防肝硬化腹水患者发生腹腔感染的约翰霍普金斯循证实践模型下中医护理策略的构建研究
  • 批准号:
    2026JJ81909
  • 项目类别:
    省市级项目
  • 资助金额:
    --
  • 批准年份:
    2026
  • 负责人:
    胡曦
  • 依托单位:
基于Deep Unrolling的高分辨近红外二区荧光分子断层成像方法研究
  • 批准号:
    12271434
  • 项目类别:
    面上项目
  • 资助金额:
    46万元
  • 批准年份:
    2022
  • 负责人:
    贺小伟
  • 依托单位:
基于深度森林(Deep Forest)模型的表面增强拉曼光谱分析方法研究
  • 批准号:
    2020A151501709
  • 项目类别:
    省市级项目
  • 资助金额:
    10.0万元
  • 批准年份:
    2020
  • 负责人:
    谢怡
  • 依托单位:
面向Deep Web的数据整合关键技术研究
  • 批准号:
    61872168
  • 项目类别:
    面上项目
  • 资助金额:
    62.0万元
  • 批准年份:
    2018
  • 负责人:
    董永权
  • 依托单位: